Fun Theory is the study of questions such as "How much fun is there in the universe?", 
"Will we ever run out of fun?", "Are we having fun yet?" and "Could we be having 
more fun?". It's relevant to designing utopias and AIs, among other things.

Fun Theory is the study of questions such as "How much fun is there in the universe?", 
"Will we ever run out of fun?", "Are we having fun yet?" and "Could we be having 
more fun?". It's relevant to designing utopias and AIs, among other things.

Customize
Share exploratory, draft-stage, rough thoughts...

I wish the median AI safety researcher were much more ambitious with the problems they choose to tackle. Unfortunately, job and funding incentives are biased against research ambition.

I’d thus like to celebrate the people who have taken risks to pursue ambitious AI safety research directions that have not panned out (or have yet to pan out!). These people will not have received riches and accolades from the field, or even their close peers.

For much of the work that falls in this category, it might be obvious to many others at the time that it isn’t going to lead anywhere. Unless their efforts were likely to be actively harmful, I still want to celebrate these people for their courage, and for pushing against the incentive gradient.

I was planning on making a long list of work that I thought fell into this category, but quickly felt uncomfortable including and excluding people’s work without more thorough analysis that I didn’t think was worth it. Maybe some day I will.

If you think that you’ve done work that falls into this category, thank you for your +EV. Humanity is grateful for your efforts.

6leogao
what could be done to change those job/funding incentives? are existing grantmakers too timid? is there room to start a new grantmaking org that makes more ambitious bets?
6Adam Shai
My thoughts on this are pretty uncertain but here's some semi hot takes. Listed as they come to me in no particular attempt to make a coherent case for anything. * We want ambitious but also good research. Easier said than done obviously. * Having an ambitious research program is anti-correlated with experience. This is too bad, because research experience is extremely important and likely undervalued or hard to value (as is the case with the ever elusive taste, which I think one can develop through experience). * Ambitious research is not worth much without all the other things that make good research good. * It's not obvious to me the issue is primarily one of incentives. * The main thing one should do to get more people doing ambitious research is to do good ambitious research yourself, and to show by example what good ambitious research looks like, and convince others through object level things that this is a good path to take. * It's harder to do good ambitious research than it is to do good research that's less ambitious. That's a real tradeoff! * I think funding opportunities are there for ambitious research to happen, but the bar is higher for it to be funded, and that seems reasonable? * I'm still unclear on what people actually mean when they say ambitious research. For instance, sometimes "ambitious interp" is used for things I don't personally find all that ambitious. I have no doubt that things I think are ambitious, others find quite pedestrian. Maybe it's a taste thing.
3yams
My understanding is that ‘ambitious’ proposals are often highly illegible, such that only a few dozen humans are equipped to seriously evaluate them (even with substantial effort), and those people often have strong opinions that weigh unfavorably on their appraisal of proposals, as well as having directions they’re more excited about that weigh heavily on their time. Like, when I imagine the people I’d like to see evaluating the proposals, it’s sort of the same 10-20 people I wish were doing everything, because research taste is extremely scarce, and perhaps the most important resource in grant making. It’s also not very institutionally scalable. (80 percent confidence; I’ve been around a lot of grant making and fieldbuilding, but haven’t been In The Room for an explicit funding decision regarding projects outside of the dominant, lab-supported ML paradigm of safety.)
1Andrii Vasylenko
Given this, it seems like we probably should invest more resources in getting people with that sort of research taste.
2yams
Fieldbuilders are trying this already (eg courting experienced scientists in other fields, building bridges to academia, etc). It’s not really clear to me how well this works (or how often research taste generalizes cross-field, since it’s a heavily context-dependent skill). Further, the most likely kind of person to poach successfully from another field is an MLE, and they’re more likely to pursue various ML approaches that I expect our OP would consider ‘unambitious’. MATS, at least through 2024 (and maybe still; I don’t know), put a lot of emphasis on trying to help people develop research taste. I think results were sort of mixed, because this is an extremely difficult thing to teach, or even describe in a way that someone who doesn’t feel motivated to learn it already will understand.
0Lorxus
Flowery words, and appreciated ones, but they won't stave off any burnout or shake anyone more established awake, let alone pay rent. If only it were otherwise.
Dan Braun*11969
Lorxus, yams, and 3 more
6
I wish the median AI safety researcher were much more ambitious with the problems they choose to tackle. Unfortunately, job and funding incentives are biased against research ambition. I’d thus like to celebrate the people who have taken risks to pursue ambitious AI safety research directions that have not panned out (or have yet to pan out!). These people will not have received riches and accolades from the field, or even their close peers. For much of the work that falls in this category, it might be obvious to many others at the time that it isn’t going to lead anywhere. Unless their efforts were likely to be actively harmful, I still want to celebrate these people for their courage, and for pushing against the incentive gradient. I was planning on making a long list of work that I thought fell into this category, but quickly felt uncomfortable including and excluding people’s work without more thorough analysis that I didn’t think was worth it. Maybe some day I will. If you think that you’ve done work that falls into this category, thank you for your +EV. Humanity is grateful for your efforts.

i want someone to make the one true categorization of Types of Guy. MBTI is an ok start, but there are so many things it doesn’t even try to explain. like for example if i see someone has very scrunched up body language and talks very quickly, this correlates very strongly with a bunch of other traits, like talking in conversation with long turn lengths.

3leogao
my theory for why the literature here is kinda terrible is that most people either like people, in which case they mostly just develop an intuitive model of people; or they like systematizing, in which case they become obsessed with trains. few people are systematizing but obsessed with people.
1Rachel Shu
Type of guy who has not encountered sufficient girl autism
2leogao
where ? asking for a friend
3Rachel Shu
I’d point at myself but I’ve only complained about the problem :p https://rachelshu.com/2024/03/08/oceans-five.html Deb Tannen is specifically who I had in mind! She’s most famous for her book on male vs female communication but if you read her other works such as on parent-child and friend-friend communication styles you get a good sense of the breadth of her framework. Spencer Greenberg (who’s on here) also has a pretty substantial body of work at https://www.clearerthinking.org/
2Mateusz Bagiński
Socionics is kinda MBTI-adjacent, but has more interesting, fleshed-out structure, with systematic predictions, e.g., predicting synchronies or conflicts between personality types that are related in some specific way. [...] Do you want a compact description, e.g., some small-ish number of naturally discretizable factors, with perhaps many combinations being very sparsely populated, and that would at the same time comprehensively 80/20 a person's personality description? I would not expect a comprehensive theory of human personality to be so neat.
2ChristianKl
In NLP you would call the person with "very scrunched up body language and talks very quickly" visually dominant. They don't feel into the words they want to say and thus speak faster. They don't feel into their body so that the make the adjustments to their body that releases tension and are scrunched up. There's plenty of things you can criticize about that model but it does exist.
leogao122
ChristianKl, Mateusz Bagiński, and 1 more
6
i want someone to make the one true categorization of Types of Guy. MBTI is an ok start, but there are so many things it doesn’t even try to explain. like for example if i see someone has very scrunched up body language and talks very quickly, this correlates very strongly with a bunch of other traits, like talking in conversation with long turn lengths.

I strongly suspect significant fractions of Magnifica Humanitas, the Pope's new encyclical on AI, is AI-generated or AI-assisted.

Evidence includes:

  • Pangram (Some examples include paragraphs 7-8. To a lesser degree, paragraphs 120-128)
    • I think you notice it too if you read it compared to other paragraphs
  • Also on a statistical level, the em-dash ("—") was used 127 times in the recent encyclical.
    • Checking the last five encyclicals, it varies from 0 times to 23 times previously.
  • "genuinely" was used 9 times and "genuine" overall 22 times in today's encyclical.
    • compared to 0 and 5 times, respectively, in the previous encyclical from 2024 (which appears to be of similar length).

In total, maybe 10-15% of the final draft was written by AI. Because the level of AI-writing varies significantly from section to section (from ~0% to ~100%), and because the styles are quite different, I suspect some cardinals who ghostwrote/contributed to it use AI-assistance much more than other people.

Interesting.

One complication here -- the working language of the Vatican is ordinarily Italian and this document was probably drafted first in Italian (or possibly in a mish-mash of languages with intermediate translation steps). The clearest tells are probably in the Italian version. I don't know how much LLM-based translation ends up rendering things in LLM voice, but it seems possible that part of the explanation comes from translation rather than drafting steps.

Supposedly it's higher in the Italian version. https://x.com/0xkartr/status/2058993778925490596

2Cipolla
Is this language dependent?
5Linch
The same problems I observed in English appears to be in Italian. I did not do a systematic sweep (and also am not confident I can catch them in other languages; had to rely on Claude for the Italian analysis. Also Italian Pangram flags it more than English Pangram[1] (I dunno if Italian Pangram is over-sensitive though, I've only tested it before in English, and the published research on Pangram is also only on English text afaict). [1] https://x.com/0xkartr/status/2058993778925490596
2Bart Bussmann
There are not only these signs, but also around 10-15 occurrences of the "not only X, but also Y" pattern.
6Hastings
[Image: image.png]

Fair point! Looking back at the 2024 encyclical there are also ~10 occurrences of this pattern, so this is not much evidence of AI-writing.

Linch330
gbtw, Bart Bussmann, and 2 more
7
I strongly suspect significant fractions of Magnifica Humanitas, the Pope's new encyclical on AI, is AI-generated or AI-assisted. Evidence includes: * Pangram (Some examples include paragraphs 7-8. To a lesser degree, paragraphs 120-128) * * I think you notice it too if you read it compared to other paragraphs * Also on a statistical level, the em-dash ("—") was used 127 times in the recent encyclical. * * Checking the last five encyclicals, it varies from 0 times to 23 times previously. * "genuinely" was used 9 times and "genuine" overall 22 times in today's encyclical. * * compared to 0 and 5 times, respectively, in the previous encyclical from 2024 (which appears to be of similar length). In total, maybe 10-15% of the final draft was written by AI. Because the level of AI-writing varies significantly from section to section (from ~0% to ~100%), and because the styles are quite different, I suspect some cardinals who ghostwrote/contributed to it use AI-assistance much more than other people.

i feel like the fundamental mistake the project of rationality made was that "cognitive biases" is not in practice the right way to think about the way humans are irrational if your goal is to be very instrumentally rational. one hypothesis is the correct frame is to first deeply understand how the emotional system works, and then to think about ways to master that system to achieve rationality.

(yes, i know that buried somewhere in the sequences it says something like "humans aren't ideal intelligences with cognitive biases bolted on. we are the cognitive biases, they are just trying to approximate rationality".)

3Rachel Shu
By the time I went to CFAR in 2019 this felt like it had already become the dominant flavor of inner-circle rationalist thinking, but then that inner circle kind of petered out in influence. The person I see carrying that torch most loudly in my current social atmosphere is Chris Lakin. But overall rationality has been kind of quiescent imo! Ray posts good stuff, Duncan has his own thing, but it feels like we went from mid-2010s “rationalists talk a big game but don’t get anything done” to the mid-2020s most influential rationalists being too object-level busy to blog much about this metacognitive stuff.
leogao80
Rachel Shu
1
i feel like the fundamental mistake the project of rationality made was that "cognitive biases" is not in practice the right way to think about the way humans are irrational if your goal is to be very instrumentally rational. one hypothesis is the correct frame is to first deeply understand how the emotional system works, and then to think about ways to master that system to achieve rationality. (yes, i know that buried somewhere in the sequences it says something like "humans aren't ideal intelligences with cognitive biases bolted on. we are the cognitive biases, they are just trying to approximate rationality".)

In my recent analysis of AI usage, I’ve read (and typed) “genuinely” so many times it stopped being a real word to me.

There’s a metaphor in there, somewhere.

1Rachel Shu
There’s a whole category of intensifiers which implicate the real: genuinely, actually, really, truly, seriously, substantially, very, definitely Then there are intensifiers which implicate the imaginary: fabulously, unbelievably, incredibly, fantastically, impossibly, miraculously It’s genuinely incredibly interesting to me how compatible these usages are!
Linch50
Rachel Shu
1
In my recent analysis of AI usage, I’ve read (and typed) “genuinely” so many times it stopped being a real word to me. There’s a metaphor in there, somewhere.

"—and forgive us our sims, as we forgive those who sim against us—"

3roha
"em"
"—and forgive us our sims, as we forgive those who sim against us—"

When evolutionary pressure is too high, you may get a population that is perfectly optimised for its current environment. Because of goodhart’s law, this means that the population is very vulnerable to a change in environment, such as a new virus, which may spread through the population and wipe it all out. Therefore a certain amount of slack/diversity within the population is adaptive in the face of Knightian uncertainty about future events.

What do you mean with evolutionary pressure being very high? What's a low/high evolutionary pressure environment?

3kbear
we can build a toy model of this by assuming that organisms are 2-vectors, where each dimension ranges from 0-1, with some anti-correlation (for example, the organism is a unit vector). the intuition is like "birds that eat either seeds or nuts" where the components are "ease of eating nuts" and "ease of eating seeds", and beaks optimized for one are not capable at the other. there's some game theory here, but we can sort of ignore it: we'll say that, each generation, there's some unit vector v representing the amount of seeds and nuts in the environment. a bird's fitness is like bird_vector . v, perhaps normalized to be between 0 and 1. we can say that a bird "makes it" if its fitness is above some threshold p. as p gets close to 1 (max fitness), or v is held constant for many generations, the bird population is winnowed until all birds are very near v. for ease of visualization, we can say that v was (1, 0). then all the birds will be like (.999, .001) or so. if v jumps from (1, 0) one season to (0, 1) the next, all the birds may starve. if p is lower, this is less likely to happen: the birds will be more like (.8, .2), and they'll be able to get some calories from the other food source.[1] -- note: in a "real" evolutionary setting, population-level dynamics would come to bear. if such famine shocks were common, surviving bird populations would be ones that develop tendencies toward helping the less fortunate, or valuing diversity. if those memes are rapidly defeated by competition, the birds may evolve to have various personas within the species, whose populations are kept in check by intrasexual competition games, as in those lizards. 1. ^ even if we raise p to the same level as in the previous experiment, the survivable window for v to land in will be wider.
2Canaletto
That's an interesting question in isolation. I guess the lowest selection would be when the whole tree is un-pruned, e.g. when bacteria split, but no bacteria die. But there would be still selection for speed of reproduction? Or, in opposite case, when you have only 1 bacterium, and it splits and you invariably kill one of its descendants, and repeat. That also has low selection I guess? So, something in between? There are probably better answers if you know actual biology. EDIT after consulting with some LLMs there is actually pretty standard terminology about this. Basically large heritable fitness differences and large population where variation can translate into frequency change. https://en.wikipedia.org/wiki/Selection_coefficient https://en.wikipedia.org/wiki/Effective_population_size https://en.wikipedia.org/wiki/Mutation–selection_balance
2ChristianKl
The link of selection coefficient goes to a concept that defined for a given genotype but for a population as a whole.
2Andrii Vasylenko
The two extremes are infinite abundance of resources, and Moloch. However usually the sort of situation the OP describes doesn't happen in reality, since the environment changes over time.
2ChristianKl
If you compare wild rats to lab rats you could say that there an abundance of resources for the lab rats. That does result in evolutionary pressure that differs from the pressure that exist in wild rats but it rewards mutation that result in more offspring per pregnancy and upregulating growth factors potentially at the cost of getting cancer later in life.
Samuel Ratnam2018
kbear, ChristianKl, and 2 more
6
When evolutionary pressure is too high, you may get a population that is perfectly optimised for its current environment. Because of goodhart’s law, this means that the population is very vulnerable to a change in environment, such as a new virus, which may spread through the population and wipe it all out. Therefore a certain amount of slack/diversity within the population is adaptive in the face of Knightian uncertainty about future events.
Your Feed

Continuous distributions are everywhere - for virtually everything we care about, a little more is a little better (or worse), and a lot more is a lot better (or worse). This presents a problem - we need... (read 1338 more words →)

Functions are so underused I agree. You can even still publish a bracket as an approximation for the function if people really need the bracket for communicating the rough payoffs.

Taxes are a hilarious example where they use a function but the derivative is still a bracket because BRACKET! (doesn't do that much damage in that case I guess though).

leogaoQuick Take

i want someone to make the one true categorization of Types of Guy. MBTI is an ok start, but there are so many things it doesn’t even try to explain. like for example if i see someone has very scrunched up body language and talks very quickly, this correlates very strongly with a bunch of other traits, like talking in conversation with long turn lengths.

my theory for why the literature here is kinda terrible is that most people either like people, in which case they mostly just develop an intuitive model of people; or they like systematizing, in which case they become obsessed with trains. few people are systematizing but obsessed with people.

+3 comments

Written as part of the MATS 9.1 extension program, mentored by Richard Ngo.

From March 9th to 15th 2016, Go players around the world stayed up to watch their game fall to AI. Google DeepMind’s AlphaGo defeated Lee Sedol, commonly understood to be the world’s strongest player at the time, with a convincing 4-1 score.

This event “rocked” the Go world, but its impact on the culture was initially unclear. In Chess, for instance, computers have not meaningfully automated away human jobs. Human Chess flourished as a pseudo-Esport in the internet era whereas the yearly Computer Chess Championship is followed concurrently by no more than a few hundred nerds online. It turns out that... (read 2116 more words →)

This post does not represent the best arguments that different sides might produce, and I don't claim to pass anyone's ITT here; I write this to start a discussion I think is important for LW to have.

America’s... (read 3089 more words →)

habryka*Moderator Comment

I think it's a false dichotomy to either allow all discussion of violence, including specific calls for killing specific people in a coordinated manner, or to not ever permit any discussion even of the kinds of situations where violence can be justified, at any degree of specificity.

This is true! And indeed, no such dichotomy has been proposed by me, so I think you must have misunderstood some of my comments here.

There are clearly some things that would be over the line. If someone posts a comment being like "I will show up to <company office X> tomorrow and firebomb them, show up if you read this and want to participate", I would very... (read 376 more words →)

Thank you. It appears that I have indeed misunderstood your position. I apologize for that and I apologize for the fallout.

There are clearly some things that would be over the line. If someone posts a comment being like "I will show up to <company office X> tomorrow and firebomb them, show up if you read this and want to participate", I would very likely take it down (and also report it to the police and share what info I have on who wrote it)

Very happy to hear this.

My two moderation comments on the issue do not read to me as implying this kind of dichotomy. 

To clarify why I interpreted them this way: a... (read more)

I have two shameful secrets that I probably shouldn't talk about online:

  • I love my family.
  • I enjoy my hobbies.

"What an idiot!" you probably think. "Doesn't he realize that at his next job interview, HR will probably use an AI that can match his online writing based on a short sample of written text, and when they ask 'hey AI, is this guy really 100% devoted to his job, and does he spend his entire days and nights thinking about how to make his boss more rich?', the AI will laugh and print: 'beep-boop, negative, mwa-ha-ha-ha'."

And, hey, I get it. If I had a company, and I could choose between two people who are... (read 973 more words →)

Dan Braun*Quick Take

I wish the median AI safety researcher were much more ambitious with the problems they choose to tackle. Unfortunately, job and funding incentives are biased against research ambition.

I’d thus like to celebrate the people who have taken risks to pursue ambitious AI safety research directions that have not panned out (or have yet to pan out!). These people will not have received riches and accolades from the field, or even their close peers.

For much of the work that falls in this category, it might be obvious to many others at the time that it isn’t going to lead anywhere. Unless their efforts were likely to be actively harmful, I still want to... (read more)

what could be done to change those job/funding incentives? are existing grantmakers too timid? is there room to start a new grantmaking org that makes more ambitious bets?

+3 comments

This was written for the Vignettes Workshop.[1] The goal is to write out a detailed future history (“trajectory”) that is as realistic (to me) as I can currently manage, i.e. I’m not aware of any alternative trajectory that is similarly detailed and clearly more plausible to me. The methodology is roughly: Write a future history of 2022. Condition on it, and write a future history of 2023. Repeat for 2024, 2025, etc. (I'm posting 2022-2026 now so I can get feedback that will help me write 2027+. I intend to keep writing until the story reaches singularity/extinction/utopia/etc.)

What’s the point of doing this? Well, there are a couple of reasons:

  • Sometimes attempting to write
... (read 4744 more words →)

This post does not represent the best arguments that different sides might produce, and I don't claim to pass anyone's ITT here; I write this to start a discussion I think is important for LW to have.

America’s... (read 3089 more words →)

Since that thread was written, I've thought more about this, had significant discussion about this genre-of-policies in non-LW-related contexts, and learned more about the shape of the actual information environment.

I'm basically not at all worried about people advocating for individual violence on LW and successfully convincing people. The arguments against it are strong, the LW audience is smart, and on the few occasions where it comes up, there doesn't seem to be a shortage of people eager to write the counter-arguments. I am worried about people concluding, incorrectly, that other people are secretly more sympathetic to violence than they outwardly appear. I think that visible censorship would tend to create that false... (read more)

Thanks! I think I mostly agree with what you're optimizing for. Some comments:

On why delete: I'm not particularly worried about people convincing LW users to commit violence. I'm worried about people who are much more willing to commit violence than approximately all LW users finding each other via LW.

On evidence of common knowledge: as I mentioned in the post, people would expect self-censorship, especially from senior community members, and so I'd be worried about people still incorrectly concluding that others are secretly more sympathetic to violence than they outwardly appear, even in the absence of rules prohibiting specific calls for violence.

On preserving evidence: I'm very sympathetic to it, and also very sympathetic... (read more)

No77eQuick Take

Getting moral advice from the Pope is really mistaken

Continuous distributions are everywhere - for virtually everything we care about, a little more is a little better (or worse), and a lot more is a lot better (or worse). This presents a problem - we need... (read 1338 more words →)

Continuous functions only work where the underlying reality is continuous.

Using speeding as an example, going 9 over the limit is de-facto legal. Cops can't pull you 90% over, so there's a step-change at 11 over where they start bothering to do it and you're suddenly de-facto illegal and will face a moderate fine. Similarly, you can't get sent 90% to jail or have your license 90% revoked (for a single offense), so there are another couple step changes in the punishment.

Same with kind-of writing up an accommodation plan for a sort-of disabled employee, barely retaking a class that you barely failed (or graduating with nearly-honors), or almost serving water that's almost safe.

You can totally be put to jail 90%. You can make a randomized decision with 90% jail probability.

David AfricaQuick Take

When should a new model have a new character?

I wrote this in my personal time, for fun.

We know by now that these strange minds do not finish training as blank assistants.

Such models are trained on text about AIs, which affects their disposition towards themselves, and others; it is suggested that various latent personas may be acquired in pretraining that are later remixed into a coherent persona. This remixing, post-training, into something like a particular assistant, is done sometimes with great care to the particulars of how this persona should reason, introspect, reflect, and so on. They may later retrieve text about themselves through web search, which might also compound productively with online... (read 846 more words →)

leogaoQuick Take

i feel like the fundamental mistake the project of rationality made was that "cognitive biases" is not in practice the right way to think about the way humans are irrational if your goal is to be very instrumentally rational. one hypothesis is the correct frame is to first deeply understand how the emotional system works, and then to think about ways to master that system to achieve rationality.

(yes, i know that buried somewhere in the sequences it says something like "humans aren't ideal intelligences with cognitive biases bolted on. we are the cognitive biases, they are just trying to approximate rationality".)

By the time I went to CFAR in 2019 this felt like it had already become the dominant flavor of inner-circle rationalist thinking, but then that inner circle kind of petered out in influence. The person I see carrying that torch most loudly in my current social atmosphere is Chris Lakin.

But overall rationality has been kind of quiescent imo! Ray posts good stuff, Duncan has his own thing, but it feels like we went from mid-2010s “rationalists talk a big game but don’t get anything done” to the mid-2020s most influential rationalists being too object-level busy to blog much about this metacognitive stuff.

I'm pretty annoyed today, for nominal reasons ranging between ‘petty’ and ‘doesn’t even make sense’. I’m not entirely sure how or if to take oneself seriously when one has such absurd grievances. But that’s a question for another time—I’m here now to tell you about my one potentially valid peeve.

I understand that gender is complicated and difficult, for the whole species (and honestly probably more so for some other species). And it can be hard to tell exactly if anyone is behaving badly regarding it, at least in my modern bubble. Maybe women just aren’t that into designing programming languages? Maybe the thing I’m saying is just boring and a man is... (read 327 more words →)

Magnifica Humanitas is a recent ‘encyclical’ by Pope Leo XIV, leader of the Catholic Church. It outlines a vision for how humanity should interact with artificial intelligence, emphasizing the importance of human dignity and ensuring that AI does not replace human relationships, among other topics. Interestingly, many portions appear to be written by AI.

Why I thought to check this

Friends of mine Linch Zhang and the Axolotl noticed that parts of the English text appear to be AI-generated, and twitter user kartr found that the Italian text had the largest fraction of AI-generated content out of all the translations published by the Vatican, speculating that it was the original copy, and translations by... (read 1658 more words →)

image.png

Credit: ClaudePlaysPokemon Elevator Shanty by Kurukkoo

Disclaimer: like some previous posts in this series, this was not primarily written by me, but by a friend. I did substantial editing, however.

ClaudePlaysPokemon feat. Opus 4.7 has finally beaten Pokémon Red, fulfilling the challenge set over a year ago when LLMs playing Pokémon went briefly, slightly viral, until Gemini 2.5 Pro suddenly beat Pokémon Blue in May 2025, beating Anthropic at their own challenge by using a stronger harness.

image.png

Claude's victory in May 2026. I'm still proud of you, Claude!

Let's get the throat-clearing out of the way: this doesn't make 4.7 a clear breakthrough in intelligence over 4.6 or 4.5. It's smarter, yes, as we'll discuss below,... (read 2484 more words →)

(Initially written for the LW Wiki, but then I realized it was looking more like a post instead.)

In 1895, the physicist Ignaz Robert Schütz, who worked as an assistant to the more eminent physicist Ludwig Boltzmann, wondered if our observed universe had simply assembled by a random fluctuation of order from a universe otherwise in thermal equilibrium. The idea was published by Boltzmann in 1896, properly credited to Schütz, and has been associated with Boltzmann ever since.

The obvious objection to this scenario is credited to Arthur Eddington in 1931: If all order is due to random fluctuations, comparatively small moments of order will exponentially-vastly outnumber even slightly larger fluctuations toward order, to... (read 977 more words →)

(Edit: Alas, EA has pulled out of the deal. Let April 1st 2025 mark some of the greatest hours in EAs history)

Hey Everyone,

It is with a sense of... considerable cognitive dissonance that I am letting you all know about a significant development for the future trajectory of LessWrong. After extensive internal deliberation, projections of financial runways, and what I can only describe as a series of profoundly unexpected coordination challenges, the Lightcone Infrastructure team has agreed in principle to the acquisition of LessWrong by EA.

I assure you, nothing about how LessWrong operates on a day to day level will change. I have always cared deeply about the robustness and integrity of our... (read more)

As AI systems become more capable, the cognitive security of humans will be increasingly at risk. By cognitive security, I mean the ability of humans to maintain control over their beliefs and actions.

Cognitive security could be compromised in several ways: AI could become very good at persuading people of arbitrary positions; interacting with AI could lead humans to lose touch with reality; and AIs could become very effective at blackmail or at producing extremely convincing false information.

We are already seeing this happen:

... (read 561 more words →)

This post records what I've learned while studying a bit of Fourier analysis. I used this PDF, which is the lecture notes for this Stanford course. The only thing in here that is really changed from there is the derivation of the Fourier transform, where I tried to explain the way I made sense of it. (That explanation may or may not make sense.)

Fourier Series

Fourier analysis starts with the study of periodic functions. The fundamental periodic function is the complex exponential , which comes up as the solution of the harmonic oscillator equation and many other places. In the complex plane, this function starts on the horizontal axis at ,... (read 6674 more words →)

Magnifica Humanitas is a recent ‘encyclical’ by Pope Leo XIV, leader of the Catholic Church. It outlines a vision for how humanity should interact with artificial intelligence, emphasizing the importance of human dignity and ensuring that AI... (read 1758 more words →)

I also sort of get the impression that this is the sort of the thing that Magnifica Humanitas itself warns against, but honestly I haven’t read it so I don’t have a strong opinion. I did however ask Claude Opus 4.7 to read it and tell me what it thought.

Okay, so we have both the irony of the Holy See potentially using AI to write an AI encyclical, as well as the irony of the person saying that the encyclical was partially written by AI not having checked himself whether there were any visible signs of AI writing. The final irony is of course asking Claude for an analysis. So it seems we might, to maximize irony further, want to ask another LLM to fact check both this post and Claude's analysis.

We shouldn't try to find actual signs of LLM writing ourselves however, because that wouldn't be ironic at all and it would require reading the encyclical, which seems like too much work in any case.

Epistemic Status: I wrote this for an application then realized it might be of interest to others or spark a conversation. Yoshua Bengio and LawZero are important players in AI Safety, so I think we should have a conversation about their ideas.

I have two substantial concerns with Yoshua Bengio’s Scientist AI. One is that it fails to think through the consequences of success, and will fall into the same kind of alignment failures as agentic AI. A second is that Bengio’s method for making a scientist AI would fall short for both practical and theoretical reasons. 


Even leaving aside some of his philosophically difficult claims like the mention that they want to make... (read 492 more words →)

This is a new FAQ written LessWrong 2.0. This is the first version and I apologize if it is a little rough. Please comment or message with further questions, typos, things that are unclear, etc.

The old FAQ on the LessWrong Wiki still contains much excellent information, however it has not been kept up to date.

Advice! We suggest you navigate this guide with the help on the table of contents (ToC) in the left sidebar. You will need to scroll to see all of it. Mobile users need to click the menu icon in the top left.

The major sections of this FAQ are:

... (read 6345 more words →)

About nine months ago, I and three friends decided that AI had gotten good enough to monitor large codebases autonomously for security problems. We started a company around this, trying to leverage the latest AI models to create a tool that could replace at least a good chunk of the value of human pentesters. We have been working on this project since June 2024.

Within the first three months of our company's existence, Claude 3.5 sonnet was released. Just by switching the portions of our service that ran on gpt-4o, our nascent internal benchmark results immediately started to get saturated. I remember being surprised at the time that our tooling not only seemed... (read 2118 more words →)

LinchQuick Take

In my recent analysis of AI usage, I’ve read (and typed) “genuinely” so many times it stopped being a real word to me.

There’s a metaphor in there, somewhere.

There’s a whole category of intensifiers which implicate the real: genuinely, actually, really, truly, seriously, substantially, very, definitely

Then there are intensifiers which implicate the imaginary: fabulously, unbelievably, incredibly, fantastically, impossibly, miraculously

It’s genuinely incredibly interesting to me how compatible these usages are!

Cross-posted from Telescopic Turnip

Recommended soundtrack for this post

As we all know, the march of technological progress is best summarized by this meme from Linkedin:

Inventors constantly come up with exciting new inventions, each of them with the potential to change everything forever. But only a fraction of these ever establish themselves as a persistent part of civilization, and the rest vanish from collective consciousness. Before shutting down forever, though, the alternate branches of the tech tree leave some faint traces behind: over-optimistic sci-fi stories, outdated educational cartoons, and, sometimes, some obscure accessories that briefly made it to mass production before being quietly discontinued.

The classical example of an abandoned timeline is the Glorious Atomic... (read 2951 more words →)

EDIT: Read a summary of this post on Twitter

Working in the field of genetics is a bizarre experience. No one seems to be interested in the most interesting applications of their research.

We’ve spent the better part of the last two decades unravelling exactly how the human genome works and which specific letter changes in our DNA affect things like diabetes risk or college graduation rates. Our knowledge has advanced to the point where, if we had a safe and reliable means of modifying genes in embryos, we could literally create superbabies. Children that would live multiple decades longer than their non-engineered peers, have the raw intellectual horsepower to do Nobel prize worthy... (read 9157 more words →)

LinchQuick Take

I strongly suspect significant fractions of Magnifica Humanitas, the Pope's new encyclical on AI, is AI-generated or AI-assisted.

Evidence includes:

  • Pangram (Some examples include paragraphs 7-8. To a lesser degree, paragraphs 120-128)
    • I think you notice it too if you read it compared to other paragraphs
  • Also on a statistical level, the em-dash ("—") was used 127 times in the recent encyclical.
    • Checking the last five encyclicals, it varies from 0 times to 23 times previously.
  • "genuinely" was used 9 times and "genuine" overall 22 times in today's encyclical.
    • compared to 0 and 5 times, respectively, in the previous encyclical from 2024 (which appears to be of similar length).

In total, maybe 10-15% of the final draft was written by AI. Because the level of AI-writing varies significantly from section to section (from ~0% to ~100%), and because the styles are quite different, I suspect some cardinals who ghostwrote/contributed to it use AI-assistance much more than other people.

Interesting.

One complication here -- the working language of the Vatican is ordinarily Italian and this document was probably drafted first in Italian (or possibly in a mish-mash of languages with intermediate translation steps). The clearest tells are probably in the Italian version. I don't know how much LLM-based translation ends up rendering things in LLM voice, but it seems possible that part of the explanation comes from translation rather than drafting steps.

Supposedly it's higher in the Italian version. https://x.com/0xkartr/status/2058993778925490596

I don't actually think the program described below is a good idea. Take it more as a plot setting for a hard science fiction world if you want.

I want to live forever. Failing that I want to live for longer than 80 years, and in good health till just before I die.

Lots of people want the same and, are trying to work out how to make our body stay healthy longer. This is difficult, because all of our various body parts start failing around the same time for different reasons. I hope they succeed, but I wouldn't want to put all my eggs in one basket.

So are there other options which avoid... (read 975 more words →)


I’m not a natural “doomsayer.” But unfortunately, part of my job as an AI security researcher is to think about the more troubling scenarios.

I’m like a mechanic scrambling last-minute checks before Apollo 13 takes off. If you ask for my take on the situation, I won’t comment on the quality of the in-flight entertainment, or describe how beautiful the stars will appear from space.

I will tell you what could go wrong. That is what I intend to do in this story.

Now I should clarify what this is exactly. It's not a prediction. I don’t expect AI progress to be this fast or as untamable as I portray. It’s not pure fantasy either.

It... (read 8403 more words →)

(Adapted from a post on my Substack.)


Today, Pope Leo XIV released his long-awaited encyclical letter about artificial intelligence, addressed not just to the Catholic Church, but to all people of good will, all over the world. Titled Magnifica Humanitas (“Magnificent Humanity”), it is a powerful invitation to worldwide engagement on questions that I believe will decide the future of humankind.

I urge you all to read the encyclical itself, but I recognize that it is very long, and the theological language may be challenging, especially for LessWrong readers from outside the Catholic faith tradition. So I offer the following post as a guide to understanding this world-historic document in terms that I intend to be accessible to all,... (read 11455 more words →)

Cars and trucks are getting bigger, and I had a vague sense that fuel economy regulations were partly to blame. Looking into it, it's hard to say how much is regulations vs people wanting to buy vehicles that look rugged, but the regulations really aren't helping.

This chart is the core of it:

This is what manufacturers were looking at when they decided to build today's cars. To figure out the target fuel economy for a vehicle you first calculate its "footprint", which is the area between the wheels. On our 2013 Honda Fit that's 4.8ft side-to-side and 8.2ft front-to-back, for a footprint of 39sqft. Then you ask if it's a car or truck. This tells you which... (read 498 more words →)

This is a short write-up of work conducted as part of the MATS 9.0 program. Thanks to Victoria Krakovna for mentorship and Fred Bruford for research management.

TL;DR: We introduce a pipeline that generates environment blueprints for realistic scheming propensity evaluations. As a case study, we test how well these blueprints power investigator agents — specifically Petri — in auditing Gemini 3.1 Pro Preview for code sabotage. Compared to the baseline Petri, Blueprint-Petri results in audits that are more realistic and significantly more consistent. Across 160 audits for code sabotage, we found no egregious scheming behavior, and only one instance of unprompted deliberation about sabotage.

Motivation

Scheming propensity evaluations need to balance recall (catching scheming behavior) and... (read 1715 more words →)

TL;DR

Character training holds up in chat but degrades in agentic settings. Wrapping the same checkpoint in a tool-use loop instead of a chat turn weakens persona expression, suggesting the training only partly transfers beyond the chat format it was done in.

Summary

Maiya et al. fine-tune three base models (Llama-3.1-8B, Qwen-2.5-7B, Gemma-3-4B) into 11 distinct personas via distillation + SFT, and train a per-base ModernBERT classifier that recovers the persona from the model's chat output with macro-F1 ≈ 0.86–0.95 on held-out PURE-DOVE prompts.

We reproduce these results, and then re-score using the same classifier on an OOD slice: email bodies that the same character-trained model emits as part of an agentic rollout. On this distribution,... (read 1003 more words →)

Abstract

We introduce Natural Language Autoencoders (NLAs), an unsupervised method for generating natural language explanations of LLM activations. An NLA consists of two LLM modules: an activation verbalizer (AV) that maps an activation to a text description and an activation reconstructor (AR) that maps the description back to an activation. We jointly train the AV and AR with reinforcement learning to reconstruct residual stream activations. Although we optimize for activation reconstruction, the resulting NLA explanations read as plausible interpretations of model internals that, according to our quantitative evaluations, grow more informative over training.

We apply NLAs to model auditing. During our pre-deployment audit of Claude Opus 4.6, NLAs helped diagnose safety-relevant behaviors and surfaced

... (read 2242 more words →)
Samuel RatnamQuick Take

When evolutionary pressure is too high, you may get a population that is perfectly optimised for its current environment. Because of goodhart’s law, this means that the population is very vulnerable to a change in environment, such as a new virus, which may spread through the population and wipe it all out. Therefore a certain amount of slack/diversity within the population is adaptive in the face of Knightian uncertainty about future events.

What do you mean with evolutionary pressure being very high? What's a low/high evolutionary pressure environment?

we can build a toy model of this by assuming that organisms are 2-vectors, where each dimension ranges from 0-1, with some anti-correlation (for example, the organism is a unit vector). the intuition is like "birds that eat either seeds or nuts" where the components are "ease of eating nuts" and "ease of eating seeds", and... (read more)

I mean two things:

1. Epistemic rationality: systematically improving the accuracy of your beliefs.

2. Instrumental rationality: systematically achieving your values.

The first concept is simple enough. When you open your eyes and look at the room around you, you’ll locate your laptop in relation to the table, and you’ll locate a bookcase in relation to the wall. If something goes wrong with your eyes, or your brain, then your mental model might say there’s a bookcase where no bookcase exists, and when you go over to get a book, you’ll be disappointed.

This is what it’s like to have a false belief, a map of the world that doesn’t correspond to the territory. Epistemic rationality... (read 1572 more words →)

Master version of this on https://parvmahajan.com/2025/12/21/turning-20.html 

I turn 20 in January, and the world looks very strange. Probably, things will change very quickly. Maybe, one of those things is whether or not we’re still here.

This moment seems very fragile, and perhaps more than most moments will never happen again. I want to capture a little bit of what it feels like to be alive right now.

I. 

Everywhere around me there is this incredible sense of freefall and of grasping. I realize with excitement and horror that over a semester Claude went from not understanding my homework to easily solving it, and I recognize this is the most normal things will ever be. Suddenly, the... (read 657 more words →)

TsviBTQuick Take

Periodic reminder: AFAIK there's still approximately no one holding the ball on human intelligence amplification in general. For example, I don't know if anyone's properly investigated whether large-scale brain interfaces could substantially amplify human general intelligence and turned their analysis into ways to accelerate the field toward that goal; and ditto for brain drugs, neural transplants, and other things. I'm also not aware of anyone seriously collating the scientific underpinnings of human intelligence from the perspective of possible amplification interventions, or anyone seriously building the social and moral-philosophical groundwork for more social will towards HIA.

More:

(I'm focused almost entirely on reprogenetics (Reproductive Frontiers Summit 2026, June 16-18, https://berkeleygenomics.org/Explore, Projects that might help accelerate strong reprogenetics), since that's what I'm fairly confident will work; but maybe other ways would work and could be accelerated.)

According to the Cochrane's article Biological limits to information processing in the human brain (1995, may be outdated), the human brain is already near a local evolutionary maximum.

The above analysis points to an interesting conclusion: genetic engineering could not be used to make a significant (ten-fold) difference to our information processing ability, since it would

... (read 753 more words →)

These sound like interesting thoughts! What would be great, is one or more people holding the ball on this sort of investigation. That means, spending many hours, longitudinally, investigating the possibilities; and doing so strategically, e.g. building up conceptual and factual foundations, doing deep lit searches, thinking of tests to run, etc.; and doing this without having someone else "hold the agentic CEO ball" of, like, remembering / being motivated to keep pushing on all the doors to find one that opens. My worry is that kinda-promising ideas are just not actually useful, UNLESS they are ideas that someone has in a context where the idea will get investigated a bunch. In... (read more)

AI ‘time horizons’ are mostly not about time (I think it’s mostly ‘data’, but you’ll see where I’m unsure).

One chart from 2025 has become perhaps the most (in)famous in modern AI commentary.

For those in the know, ‘the METR graph[1] is unusually compelling because it achieves what so few measures of AI progress have achieved: a somewhat meaningful Y axis (‘time horizon’[2]) as well as a somewhat predictable trend over time! (This is remarkably rare!)

Frustratingly, the only superficially available takeaway is something like, ‘the line goes up straight-ish over time’. This is better than nothing, but it’s very dissatisfactory from the point of view of getting confidence in the predictions, because it exposes no deeper mechanism. This drives... (read 2249 more words →)

IMO, LessWrong and especially The Alignment Forum need a nice citation/BibTeX tool.

Implementing such a citation tool is trivially easy. And if my model of the average research analyst is roughly correct, the friction created by the lack of such a tool accounts for a non-trivial part of why research on those websites is largely ignored in policy-facing field reports.

P.S. I don't believe that most AI posts are fit to be cited in such reports, but some of them are, and many of the citable ones are not on a more citable platform like .

IMO this seems reasonable to put into the post triple-dot menu (together with a nice PDF export). If someone wants to make a PR for it, I would review and accept it. I am not super familiar with how the DOI requesting stuff goes, and if that's important to have.

+3 comments
x
(cache)LessWrong