Jonathan Stray

25K posts

Jonathan Stray banner
Jonathan Stray

Jonathan Stray

@jonathanstray

Knowing things is a solved problem. Getting along is not. Doing the science to ensure AI results in less war, not more @CHAI_Berkeley.

Berkeley Katılım Mayıs 2008
2K Takip Edilen10.6K Takipçiler
Sabitlenmiş Tweet
Jonathan Stray
Jonathan Stray@jonathanstray·
You may remember we did a 10,000 person experiment testing AI-powered healthier feed algorithms. Well, now we’ve built it as a product for BlueSky/ATProto (because X doesn't support 3rd party algos). We are currently recruiting pre-release users to test it out and give us feedback. Want to try it?survey.qualtrics.com/jfe/form/SV_01…
Jonathan Stray tweet media
English
0
4
26
4.9K
Jonathan Stray
Jonathan Stray@jonathanstray·
@MeganTStevenson You may be right (and I love your scholarship BTW and I respect your opinion). Still doesn't let humanities off the hook.
English
1
0
3
131
Megan Stevenson
Megan Stevenson@MeganTStevenson·
It's interesting that the humanities are getting the bad rap. IMO the majority of empirical economics is quasi-fraudulent, in that the researcher is bending the data to tell the most publishable story possible.
Brendan Nyhan (@BrendanNyhan on 🟦☁️)@BrendanNyhan

All can be true: 1. Universities have been targeted by a decades-long bad faith campaign 2. Status quo is vastly more normal than craziest stuff on Fox/social media, esp. outside elite schools 3. Parts of academia are in rough shape Despite 1 & 2, we can and should address 3.

English
11
14
122
15.1K
Jonathan Stray
Jonathan Stray@jonathanstray·
This is my entry into the "what's wrong with humanities scholarship" discourse. I'm explaining how and why I think the politics of these fields are holding them back. I hope this is thoughtful and well-considered.
Jonathan Stray@jonathanstray

There's a longer answer to this question, of course. Every field has some mix of normative and descriptive work. This can get unbalanced. I've been arguing for computer scientists to take clearer normative stances, for example. From the outside, what I would say is in the last few decades the humanities seem to me to have become dominated by, let's call it the critical turn. The emphasis on race/class/gender oppression, and more generally critique instead of description, is diagnostic. These fields also give off a more insular vibe, at least to me; for example I don't see a lot of work seriously grappling with how economics actually works (IMO Graber's "Debt" was unenlightening), and some major developments in psychology seem ignored too. Certainly the politics of humanities scholars have collapsed; there are many ways to see this, but there's lots and lots of data and it all says the same thing: there's some sort of leftist uniformity dominating all of these fields. Leftist politics are not intrinsically bad of course, but it does have its dysfunctions, and these are amplified by uniformity of thought. At the same time, there are definitely right-wing activists attacking the humanities. Yet many external criticisms cannot simply be rejected as insincere or uninformed. So this is the current situation as I see it: multiple related humanistic fields have become overbalanced with a fairly narrow kind of prescriptivist approach. In that context, more descriptive work can be neglected, the field starts to feel insular, and people are going to be more suspicious of political claims coming from affiliated scholars. And unfortunately, I think at least some of that suspicion is justified. If nothing else, this is a parsimonious explanation for the collapse of trust in humanities scholarship. You may not agree with me, of course. This debate has been endlessly rehashed, and there is a very active Discourse around all of this, as I am sure you know. But hopefully this at least gives you a sense of how I, personally, interpret the state of affairs as suggesting that normative restraint would be corrective.

English
1
0
1
380
Jonathan Stray
Jonathan Stray@jonathanstray·
How does the GreenEarth feed (available over on BlueSky) determine "constructiveness"? We use the Google Perspective API but not the classic "toxicity" score. Instead we combine many of their "bridging" attributes. The weights are from our previous 10,000 person experiment. Find our feed here: greenearth.social/p/welcome-to-t… Refs: #L58" target="_blank" rel="nofollow noopener">github.com/greenearth-soc… medium.com/jigsaw/announc… rankingchallenge.substack.com/p/its-possible…
Jonathan Stray tweet media
English
0
0
2
170
Jonathan Stray
Jonathan Stray@jonathanstray·
I'm not so sure community notes aren't impactful. I suspect they work better than human notes ever did -- I still haven't seen a direct comparison. Anyway, they can just publish elsewhere, lots of friendlier venues. (This is also warning about the insular politics of tech criticism, I think.)
English
1
0
1
26
Robyn Caplan
Robyn Caplan@robyncaplan·
@jonathanstray @jaybaxter I wish TPRC hadn't been so discouraging – the other reviews were basically 'community notes aren't impactful, so this is moot.' This was really a student-led paper (an undergrad). If I had all the time in the world, I'd try and make sure this was in a publishable spot.
English
1
0
0
25
Robyn Caplan
Robyn Caplan@robyncaplan·
Btw last month I got a rejection from TPRC on a paper about the community notes algorithm (co-authored with a student) and one review was literally just “be careful what you wish for.” What do we think they meant by this?
English
4
0
5
968
Robyn Caplan
Robyn Caplan@robyncaplan·
@jonathanstray @jaybaxter I had not seen this. What was really interesting about my student's findings was that many "fact-checking" norms and practices actually persist in community-based models.
English
2
0
0
34
Jonathan Stray
Jonathan Stray@jonathanstray·
GreenEarth (@greenearth-social.bsky.social) will be the first feed to explain the algorithm, tell you why you got each post, and let you adjust all the parameters. Here's an early version running on our staging server. We're ~2 weeks away from public beta.
Jonathan Stray tweet media
English
1
0
3
285
Jay Baxter
Jay Baxter@jaybaxter·
@robyncaplan @jonathanstray oof. sadly we got a similarly weird review on this topic recently too... made me think it'd be great to try to improve the peer review process with some ideas from the community notes algo ha
English
1
0
1
41
Jonathan Stray
Jonathan Stray@jonathanstray·
There's a longer answer to this question, of course. Every field has some mix of normative and descriptive work. This can get unbalanced. I've been arguing for computer scientists to take clearer normative stances, for example. From the outside, what I would say is in the last few decades the humanities seem to me to have become dominated by, let's call it the critical turn. The emphasis on race/class/gender oppression, and more generally critique instead of description, is diagnostic. These fields also give off a more insular vibe, at least to me; for example I don't see a lot of work seriously grappling with how economics actually works (IMO Graber's "Debt" was unenlightening), and some major developments in psychology seem ignored too. Certainly the politics of humanities scholars have collapsed; there are many ways to see this, but there's lots and lots of data and it all says the same thing: there's some sort of leftist uniformity dominating all of these fields. Leftist politics are not intrinsically bad of course, but it does have its dysfunctions, and these are amplified by uniformity of thought. At the same time, there are definitely right-wing activists attacking the humanities. Yet many external criticisms cannot simply be rejected as insincere or uninformed. So this is the current situation as I see it: multiple related humanistic fields have become overbalanced with a fairly narrow kind of prescriptivist approach. In that context, more descriptive work can be neglected, the field starts to feel insular, and people are going to be more suspicious of political claims coming from affiliated scholars. And unfortunately, I think at least some of that suspicion is justified. If nothing else, this is a parsimonious explanation for the collapse of trust in humanities scholarship. You may not agree with me, of course. This debate has been endlessly rehashed, and there is a very active Discourse around all of this, as I am sure you know. But hopefully this at least gives you a sense of how I, personally, interpret the state of affairs as suggesting that normative restraint would be corrective.
English
1
0
3
449
Gabriella "Biella" Coleman
Gabriella "Biella" Coleman@BiellaColeman·
@jonathanstray Legal scholarship and its normative recs *are* inherently political, not partisan, but oriented toward legal regimes already shaped by politics. My objection is the double standard: this work is seen as legitimate, while political orientation in other fields is cast as suspect.
English
2
0
4
163
Gabriella "Biella" Coleman
Gabriella "Biella" Coleman@BiellaColeman·
Explain to me like I’m a five year old: why can legal scholars and political scientists be overtly political with their policy work to take one example and anthropologists or sociologists can’t?
English
6
3
29
2.4K
Andy Hall
Andy Hall@ahall_research·
Across prediction markets, politics and similar topics seem to have the highest rates of resolution disputes.
Andy Hall tweet media
Andy Hall@ahall_research

Our new research: to start building towards political superintelligence by using AI to help us govern better, we created an AI that reads contracts and predicts ambiguities that will lead to disputes. We tested it on 10,000 Kalshi and Polymarket resolution rules, and it works quite well! Just by reading the resolution rules, our tool is able to assess which prediction-market contracts are likely to lead to resolution disputes, far better than random chance. So much of governance is writing good rules. But writing good rules is hard. For centuries, we've relied on human experts to try to write rules that anticipate as many confusions, ambiguities, and loopholes as possible. Inevitably, we make mistakes---from the famous contract law example of "the two ships Peerless" that I explore in today's piece, to unclear content moderation policies, ambiguous legislation, or the recent dustup around the US-Iran cease fire agreement wording, this is a perennial challenge. Can AI help us do it better? At Free Systems, a big part of our vision is figuring out how AI can improve how we govern, so we were eager to put this to the test. We collected a sample of 10,000 prediction-market contracts with their stated resolution rules, along with info on which ones ended up disputed. Then, we worked with our buddy Claude to develop a 10-point rubric for contract clarity, covering elements like whether the key question is well defined, whether the entities are identifiable, whether the time window is clearly specified, and so on. We had an LLM grader apply the rubric to the contracts, then built a simple machine-learning model that uses the 10-dimensional rubric score to predict subsequent disputes. The resulting scores allow us to provide overall grades to prediction-market contracts which reflect how clearly written they are and how likely they are to fall into dispute later. The contracts we grade "CCC" are 3.4x more likely to fall into dispute than the ones we grade "A." There's a lot more work to do here---we need to make sure our predictions hold in a truly out-of-sample test where we grade contracts now and see if they get disputed in the future, which we'll be working on next---and we need to expand this beyond prediction market contracts as well. But we're super excited about this direction. Tools like this will help us to identify contract ambiguities before they become disputes, allowing us to write better rules, improve governance, and eventually, get to political superintelligence. There's lots more info in our write-up, here: freesystems.substack.com/p/superintelli… Joint work with @elliotjpaschal

English
3
3
27
3.6K
Robyn Caplan
Robyn Caplan@robyncaplan·
@jonathanstray Never really got the vibe of BlueSky. I don't like a chronological feed and all of the various 'algo options' weren't intuitive unless you're using it everyday. Anyhoo...hope you wanted all this detail. Miss you @jonathanstray!
English
1
0
0
38
Robyn Caplan
Robyn Caplan@robyncaplan·
I've been taking all my tweets to my group chats over the last couple of years. Perhaps it is time to come back.
English
2
0
7
582
Jonathan Stray retweetledi
Andy Hall
Andy Hall@ahall_research·
Our new research: to start building towards political superintelligence by using AI to help us govern better, we created an AI that reads contracts and predicts ambiguities that will lead to disputes. We tested it on 10,000 Kalshi and Polymarket resolution rules, and it works quite well! Just by reading the resolution rules, our tool is able to assess which prediction-market contracts are likely to lead to resolution disputes, far better than random chance. So much of governance is writing good rules. But writing good rules is hard. For centuries, we've relied on human experts to try to write rules that anticipate as many confusions, ambiguities, and loopholes as possible. Inevitably, we make mistakes---from the famous contract law example of "the two ships Peerless" that I explore in today's piece, to unclear content moderation policies, ambiguous legislation, or the recent dustup around the US-Iran cease fire agreement wording, this is a perennial challenge. Can AI help us do it better? At Free Systems, a big part of our vision is figuring out how AI can improve how we govern, so we were eager to put this to the test. We collected a sample of 10,000 prediction-market contracts with their stated resolution rules, along with info on which ones ended up disputed. Then, we worked with our buddy Claude to develop a 10-point rubric for contract clarity, covering elements like whether the key question is well defined, whether the entities are identifiable, whether the time window is clearly specified, and so on. We had an LLM grader apply the rubric to the contracts, then built a simple machine-learning model that uses the 10-dimensional rubric score to predict subsequent disputes. The resulting scores allow us to provide overall grades to prediction-market contracts which reflect how clearly written they are and how likely they are to fall into dispute later. The contracts we grade "CCC" are 3.4x more likely to fall into dispute than the ones we grade "A." There's a lot more work to do here---we need to make sure our predictions hold in a truly out-of-sample test where we grade contracts now and see if they get disputed in the future, which we'll be working on next---and we need to expand this beyond prediction market contracts as well. But we're super excited about this direction. Tools like this will help us to identify contract ambiguities before they become disputes, allowing us to write better rules, improve governance, and eventually, get to political superintelligence. There's lots more info in our write-up, here: freesystems.substack.com/p/superintelli… Joint work with @elliotjpaschal
Andy Hall tweet media
English
15
32
179
31.4K
Andy Hall
Andy Hall@ahall_research·
The models lean left in large part because they are trained on huge amounts of text from the open internet that was disproportionately created by people with left-wing views. You can't reverse this lean through ex post prompting alone: it requires effort and money. Gemini is the most neutral probably because Google poured immense amounts of resources into specialized training to "undo" this lean. SpaceXAI is a much leaner company, and hasn't put those same level of resources into changing the lean of the model despite their public commitment to that cause.
Matthew Yglesias@mattyglesias

To follow up on this — feel free to totally ignore the results for ChatGPT and Anthropic and dismiss them as biased. Grok which is explicitly designed to not have left-wing bias still leans left! Is that because most writers are on the left? Maybe but doesn’t Grok know that?

English
10
6
69
11.5K
Jonathan Stray
Jonathan Stray@jonathanstray·
I sometimes hear concern that bridging-based ranking will homogenize discussions, flattening minority perspectives. Many answers, but the first one is: not all minority perspectives are worth preserving. I'm sure you can think of a few. We need a criterion besides "minority."
English
0
0
5
337
Jonathan Stray
Jonathan Stray@jonathanstray·
@TylerAlterman @samhbarton I could never see the horizon in NYC. Here, I see the ocean from my bedroom and ride my bicycle up into the hills behind my backyard, year round. I love this place.
English
0
0
0
30
Tyler is finishing a book, slow to reply
@samhbarton Positivity without wholesomeness + goldminer mentality (ppl here not because of a deep love of the place but because they want to get a narrow thing from it; if that could get that thing elsewhere better, they would leave)
English
4
0
34
1.8K
Jonathan Stray
Jonathan Stray@jonathanstray·
Another attempt. This isn't anywhere near our final feature set btw -- we're adding prompt-based scorers and other things -- and yet it's still more control than any social media user has ever had. Which is absurd.
Jonathan Stray tweet media
English
1
0
2
170
Jonathan Stray
Jonathan Stray@jonathanstray·
GreenEarth will be the first fully transparent, controllable, open-source feed. Does this diagram help you understand how the core algorithm works? Try the pre-release version now! It makes BlueSky useful instead of annoying! survey.qualtrics.com/jfe6/form/SV_0…
Jonathan Stray tweet media
English
2
2
10
1K