LLM-generated code that seemingly went without much review or design is always such an interesting dive into just how bloated you can make code. In this repository, multiple files are close to 10K LOC, one file contains a switch statement that has so many case statements it spans more than 1000 lines, and lots of other fun stuff.
I guess it depends on the model you're trying to use, but seems most of them prefer smaller codebases, they work a lot better with less code, which kind of makes sense. With that in mind, I'd probably aim for something way smaller to bootstrap a self-improving agent. Then I'd use this "Prime Agent" as an example to my self-improving agent for what it should not evolve to.
Literally says who. Swe are the least credible "engineers". "It depends" "No you cant track us" "No we have no credentials system other than big company shill certs"
As models get stronger, huge harnesses may become less useful. An overly opinionated harness could even constrain the model’s reasoning instead of improving it.
It'll be really interesting when they run RL training on the harness self-improvement loop. I've tried using LLMs for harness engineering, but it often creates too much bloat that weighs things down in the end. Guessing it's just not something the models are tuned to do by default.
Curious if anyone's tried using RL for harness engineering? I think we're still pretty far away from the optimal harness, especially when it comes to long-context memory management.
I vouch for this story, personally. I love it. I consider it a classic among science fiction novellas.
It does include some brutal sadomasochism with graphic torture, that's true, and that's why it isn't better known. If that part bothers you, which is fair, I would skim past it. The story overall is quite worth it
Tbh the torture scenes are not nearly as disturbing as the underage incest. It's like "Saw" vs. a pedo flick from the 70s. Especially weird since the latter is portrayed like some kind of innocent ideal.
It's frustrating because apart from those (awful) aspects the story is quite thoughtful and compelling. I always recommend people skip the final chapter since the resulting existential cliffhanger is far better than the bizarre last-minute swerve into gratuitous sex abuse.
You shouldn't link that without mentioning that it is a particularly intense S&M/rape/incest fetish piece, and that the squick content is entirely integral to the story. I'm generally dubious of "trigger warnings" but if ever there was something that needed content tags-- this is it.
If you imagine it being written for alt.sex.stories.moderated but somehow failing to be erotic by virtue of being too explicit, and being a few orders of magnitude better writing for that venue... then you wouldn't be too far off. [Started making a silly example of how some sex story could fail like that and then realized I was litterally describing a scene from prime intellect].
It's also arguably the origin of a lot of the mentally ill ai-safety hysteria. Arguably an enjoyable romp for those who understand that it's fiction, but it seems a lot of people cannot.
For everyone downvoting: It's literally a story about the creation of mankind's first artificial general intelligence, Prime Intellect, and the consequences of that discovery.
I feel that more warning is needed for this book. Everything you’ve stated is true, but the book’s first chapter contains some of the most disturbing depictions of stuff I don’t know I can type here on HN. Additionally, the last chapter put a really bad taste in my mouth.
That said, the stuff that deals with AI and its implications on the universe was great and more relevant than ever. The way PI works is pretty similar to how we use subagents to tackle large repos!
Not just the first AGI, but the book deals with a (very) hard-takeoff scenario of the singularity. The AI, Prime Intellect starts recursively increasing its capabilities and things rapidly spin out of control. This is actually a small part of the book, most of the book is about what happens after this - Its a fantastic read, yes as other posters have said there is some extremely graphic sex and violence in this book, but it is not gratuitous. I highly reccommend.
I built one of these RLM harnesses and a local MCP server along with logging, memories, and project rules based on directories. It worked great for a while but the foundational models have largely caught up to the point where they don't need this harness anymore. At least for my use cases. I can basically just store context in .md in the directories we work out of together and accomplish what I need.
I went the skills route - skill improvement skill and a rule to use it always. Basically it's directed that if anything causes more than a hop of thinking - failure - try something else it should flag that it needs to learn it as a skill so it never does it other than first shot again or if an existing skill fails improve it after it solves whatever problem. it then syncs the files to a shared location and updates the version and also pulls new skills. this let's a team use it or you have multiple workstations.
The core idea of the RLM paper is to make a regular LLM act more like a coding agent - offload context to something external that needs to be explicitly queried instead of filling up valuable context. The "recursion" part of the paper really only wins because they use a top-tier model for the root agent, and cheaper models for the sub-agents.
Prime Agent took the RLM idea (which is really just an academic view on how coding agents have always worked) and then added this "continual harness" idea. This part isn't super well described in the blog post, but includes some message passing between the agents, and the ability to share code.
Overall I chalk it up as neat, but not revolutionary. Another version of what most of these systems are already doing.
It’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers.
There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.
LLM-generated code that seemingly went without much review or design is always such an interesting dive into just how bloated you can make code. In this repository, multiple files are close to 10K LOC, one file contains a switch statement that has so many case statements it spans more than 1000 lines, and lots of other fun stuff.
I guess it depends on the model you're trying to use, but seems most of them prefer smaller codebases, they work a lot better with less code, which kind of makes sense. With that in mind, I'd probably aim for something way smaller to bootstrap a self-improving agent. Then I'd use this "Prime Agent" as an example to my self-improving agent for what it should not evolve to.
here is a take on a minimal agent ("smol")
https://github.com/smol-env/smoleasier to add and customize stuff when you start from a small base
think of it as your starter dough
> one file contains a switch statement that has so many case statements it spans more than 1000 lines
Probably best to leave YandereDev's code out of the training data.
Interesting, so they shipped slop? Succumbed to their own AI psychosis?
Bash is all you need.
https://minimal-agent.com/
Literally says who. Swe are the least credible "engineers". "It depends" "No you cant track us" "No we have no credentials system other than big company shill certs"
this seems like its going to rip through tokens like crazy
self improvement is not a new idea but at current economics its not feasible
As models get stronger, huge harnesses may become less useful. An overly opinionated harness could even constrain the model’s reasoning instead of improving it.
It'll be really interesting when they run RL training on the harness self-improvement loop. I've tried using LLMs for harness engineering, but it often creates too much bloat that weighs things down in the end. Guessing it's just not something the models are tuned to do by default.
Curious if anyone's tried using RL for harness engineering? I think we're still pretty far away from the optimal harness, especially when it comes to long-context memory management.
What policy would you use?
https://localroger.com/prime-intellect/mopiidx.html
I vouch for this story, personally. I love it. I consider it a classic among science fiction novellas.
It does include some brutal sadomasochism with graphic torture, that's true, and that's why it isn't better known. If that part bothers you, which is fair, I would skim past it. The story overall is quite worth it
Tbh the torture scenes are not nearly as disturbing as the underage incest. It's like "Saw" vs. a pedo flick from the 70s. Especially weird since the latter is portrayed like some kind of innocent ideal.
It's frustrating because apart from those (awful) aspects the story is quite thoughtful and compelling. I always recommend people skip the final chapter since the resulting existential cliffhanger is far better than the bizarre last-minute swerve into gratuitous sex abuse.
You shouldn't link that without mentioning that it is a particularly intense S&M/rape/incest fetish piece, and that the squick content is entirely integral to the story. I'm generally dubious of "trigger warnings" but if ever there was something that needed content tags-- this is it.
If you imagine it being written for alt.sex.stories.moderated but somehow failing to be erotic by virtue of being too explicit, and being a few orders of magnitude better writing for that venue... then you wouldn't be too far off. [Started making a silly example of how some sex story could fail like that and then realized I was litterally describing a scene from prime intellect].
It's also arguably the origin of a lot of the mentally ill ai-safety hysteria. Arguably an enjoyable romp for those who understand that it's fiction, but it seems a lot of people cannot.
For everyone downvoting: It's literally a story about the creation of mankind's first artificial general intelligence, Prime Intellect, and the consequences of that discovery.
I feel that more warning is needed for this book. Everything you’ve stated is true, but the book’s first chapter contains some of the most disturbing depictions of stuff I don’t know I can type here on HN. Additionally, the last chapter put a really bad taste in my mouth.
That said, the stuff that deals with AI and its implications on the universe was great and more relevant than ever. The way PI works is pretty similar to how we use subagents to tackle large repos!
yea the beginning is pretty gory / horrific. it almost made me stop reading but you can't deny the world building is.. unique
Yeah, there's no question this one's over-the-top when it comes to trying to gross out the reader, but I'd like to think people can overlook that.
Not just the first AGI, but the book deals with a (very) hard-takeoff scenario of the singularity. The AI, Prime Intellect starts recursively increasing its capabilities and things rapidly spin out of control. This is actually a small part of the book, most of the book is about what happens after this - Its a fantastic read, yes as other posters have said there is some extremely graphic sex and violence in this book, but it is not gratuitous. I highly reccommend.
I built one of these RLM harnesses and a local MCP server along with logging, memories, and project rules based on directories. It worked great for a while but the foundational models have largely caught up to the point where they don't need this harness anymore. At least for my use cases. I can basically just store context in .md in the directories we work out of together and accomplish what I need.
I went the skills route - skill improvement skill and a rule to use it always. Basically it's directed that if anything causes more than a hop of thinking - failure - try something else it should flag that it needs to learn it as a skill so it never does it other than first shot again or if an existing skill fails improve it after it solves whatever problem. it then syncs the files to a shared location and updates the version and also pulls new skills. this let's a team use it or you have multiple workstations.
The core idea of the RLM paper is to make a regular LLM act more like a coding agent - offload context to something external that needs to be explicitly queried instead of filling up valuable context. The "recursion" part of the paper really only wins because they use a top-tier model for the root agent, and cheaper models for the sub-agents.
Prime Agent took the RLM idea (which is really just an academic view on how coding agents have always worked) and then added this "continual harness" idea. This part isn't super well described in the blog post, but includes some message passing between the agents, and the ability to share code.
Overall I chalk it up as neat, but not revolutionary. Another version of what most of these systems are already doing.
A write up on RLM (Recursive Language Models) by one of the authors of the RLM paper:
https://alexzhang13.github.io/blog/2025/rlm/
We (no affiliation) have a very-alpha prototype hosted service for this, knock-on-wood available shortly: https://amdahl.sh/
Targeting a way to let people pay a few $ and spin up prime-agent experiments. Email hello@amdahl.sh if you have questions/feature requests.
It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010.
I am curious - how does it fare for other benchmarks, or everyday programming?
PrimeIntelect is not on official ARC-AGI-3 leaderboard: https://arcprize.org/leaderboard
Good to know!
Is it that it wasn't accepted yet, or are there issues with how it was run?
It’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers.
There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.
Might actually try this