Blog
Can AI Improve Its Own Ability to Improve? A New Survey Grades the Evidence
Can AI improve its own ability to improve? A new 57-page survey, co-authored by Actava.ai co-founder Weiran Yao and researchers from Tsinghua, Peking University, CMU, UC Berkeley, and other teams, reviews 404 works and offers a careful answer. Sometimes, in limited settings. The paper separates 3 claims that usually get blurred together (task gain, retention, and improver gain) and names feedback quality, interacting updates, and independent validation as the constraints on every improvement loop. This post walks through the framework, grades our own CURA training loop against it, and gives healthcare leaders 3 questions for any vendor that says "self-improving."
By Weiran Yao
Can an AI system improve its own ability to improve?
That question has a name, recursive self-improvement (RSI), and as of September 30 it has a 57-page answer.
Recursive Self-Improvement in AI: A Survey is now up as a preprint. Actava.ai co-founder and Chief AI Officer Weiran Yao is one of its 32 authors, alongside researchers from Tsinghua, Peking University, CMU, UC Berkeley, NTU, NUS, PolyU, and other academic and industry teams.
Actava.ai advisor Caiming Xiong is on the author list too, and Qingsong Wen is the corresponding author.
The short answer they landed on is "sometimes, in limited settings." The long answer is more useful, especially if you run healthcare operations and someone is selling you agents that learn.
A loop that repeats is cheap. A loop that compounds needs proof.
Most AI systems already run some kind of loop. Do the work, collect feedback, adjust, repeat.
The survey draws a hard line through that picture. Running the loop again tomorrow shows that you own a loop. Recursive improvement asks for evidence that each change the system keeps makes the next round of improvement more effective.
The paper's mechanism diagram has 2 loops in it, and the gap between them is the whole argument.
Loop 1
The system improvesThe task system (model, code, memory, tools) does the work. Its traces, outcomes, and failures feed candidate changes. Validation decides what stays. The agent gets better at the job.
External anchors
Task goals, protected evaluation, and resource budgets sit outside both loops. The system can't edit them.
Here's the healthcare version. A utilization management team that clears more cases this month than last has improved. A team whose training program turns out stronger reviewers, faster, every year has improved how it improves.
Both are good. The second one compounds.
The survey packs the distinction into 8 words: "Recursion is a dependency; improvement is an observation." You can wire a system to feed on its own output in an afternoon. Whether it got better is something you have to go measure.
3 claims that usually get blurred into one
When someone calls an agent "self-improving," the survey says they could mean any of 3 things. Each one needs its own test.
| Claim | The question | The test | Inside a health plan |
|---|---|---|---|
| Task gain | Is the system better at the work? | Score it on sealed tasks the loop never touched. | Prior authorization accuracy on cases held back from tuning. |
| Retention | Did it keep what it already knew? | Re-run earlier task groups after every update. | The appeals fix left eligibility checks intact. |
| Improver gain | Is the process that makes updates better now? | Same starting agent, same budget, old improver against new. | This quarter's tuning cycle produces a stronger successor than last quarter's would have. |
Our read is that plenty of "self-improving" pitches stop at the first row, with the system grading its own homework.
The survey's proposed protocol closes those gaps one at a time. Test data stays sealed, and its scores never flow back into what gets accepted. Every comparison runs under the same budget, with a frozen-improver control beside it. The cost ledger counts the retries, the rejected candidates, and the failed attempts (the part of the bill that tends to fall off the slide).
One question from the paper belongs on every procurement checklist. What survives a reset? A revised answer survives nothing. A stored memory, a retrained weight, or a rewritten improvement procedure carries forward, and each one calls for a different test.
What the evidence shows so far
The authors went through the strongest published systems and reported the results with the caveats attached.
The Darwin Gödel Machine, a coding agent that rewrites its own agent code, climbed from 20% to 50% on a 200-task software benchmark subset over 80 iterations. That's real progress. The survey also points out that those scores include tasks used during the search, which keeps them short of a fully independent test.
Another system let a program rewrite its own improvement routine. With a stronger base model it got better in the early rounds. With weaker models it sometimes got worse.
The idea is 23 years old in its formal version (the Gödel Machine, 2003). Foundation models made parts of it practical through generated training data, program revision, and automated experiments.
The paper's summary is careful. Systems can keep useful changes. In limited settings, they improve how they produce the next update. The evidence stops short of open-ended or reliably accelerating self-improvement, and none of it shows a loop running free of goals and feedback supplied from outside.
If a vendor tells you their agents improve themselves indefinitely, there's now a 404-reference bibliography that would like a word.
3 constraints, and healthcare sits inside all of them
The survey names 3 things that limit every improvement loop. Feedback quality, interacting updates, and independent validation.
Healthcare operations hit all 3 before lunch. In our view, that makes healthcare the hardest and most useful place to do this work.
Feedback quality
A model that grades its own work can drift. The survey warns that internal scores can keep rising while performance on the real task stays flat.
The fix is information the loop can't produce by itself. Executable tests, fresh observations, expert review. Healthcare has this in unusual supply. A nurse reviewer overturns an agent's determination. A medical director rewrites a rationale. A claim pays or it denies.
Each of those is a labeled correction from someone accountable for the outcome. It's the scarcest input in the loop, and it usually evaporates in an inbox.
Interacting updates
Changes collide. A new prompt shifts what gets retrieved. A revised tool breaks a skill that passed last month. In healthcare, add a payer policy update that lands on a Tuesday and quietly invalidates a workflow that was fine on Monday.
The survey's answer is unglamorous. Protect the capabilities you already have, set a regression tolerance before the update ships, and keep a rollback you've tested.
Independent validation
The grader has to sit outside what the loop can edit. Sealed tasks, a protected evaluator, and accounting the system can't touch.
This one is close to home. We built Actava.ai χ-BENCH to give healthcare an independent measuring stick for long-horizon, policy-dense workflows. Across 30 agent and model configurations, the best one resolved 28% of tasks. None cleared 20% when the bar was passing 3 runs out of 3.
That second number makes the survey's point from another angle. One good run tells you little. The same result 3 times in a row is what production asks for. (A benchmark score is evidence, and production still gets the final vote.)
Frank Wang put the same idea in a 10-word title in September, The Flywheel Only Spins Where You Can Check the Work.
How we'd grade our own work
It'd be easy to cite this survey and then exempt ourselves. Here's Actava.ai CURA under the same 3 claims.
We trained CURA 1T through what our paper calls a human-gated self-evolution loop. Each round, a training agent picks a target capability, trains the model, evaluates benchmark trajectories, and reworks the data mix from the failures it finds. Humans gate each round.
Task gain. CURA improved over its base model on every healthcare benchmark in the paper.
Retention. It stayed competitive on out-of-domain reasoning and agentic benchmarks, the skills it had before we specialized it.
Improver gain. Untested. We make no claim there.
In the survey's vocabulary, that human gate is an external anchor. We put it there on purpose. In a domain where a wrong determination delays someone's care, the decision to accept an update belongs to a person with a name.
Where Actava.ai sits
Actava.ai turns production workflows into intelligence customers own. The survey describes the loop in research terms. Here's the same loop in ours.
Actava.ai KORA runs the work and captures the experience (traces, outcomes, and expert corrections) inside the boundary the customer authorizes.
Actava.ai χ-BENCH turns execution failures and expert corrections into simulations and evaluations, which gives every proposed change something independent to answer to.
Verified training data from that process improves the agent harness and specialized models. Actava.ai CURA demonstrates the post-training capability, with a path toward customer-private models trained on an organization's own work.
Actava.ai Compliance supports the policies, permissions, and audit evidence around each release, inside the customer's existing governance program.
Customer-private learning stays inside its authorized boundary. The intelligence that compounds belongs to the customer.
3 questions for any vendor that says "self-improving"
You can borrow the survey's framework without reading all 57 pages.
- What survives a reset, and can you show it to me?
- Who grades each update, and can the system edit its own grader?
- What did the rejected attempts cost, and how fast can you roll back?
If the claim is that the improvement process itself keeps getting better, ask for the control. Same starting agent, same budget, old improver against new.
The open problem
Building reliable improvement loops remains an open research problem. The survey says so in its abstract, and Weiran signed it.
Our bet is that healthcare is where the problem gets solved first, because the ingredients are already on the floor. Real workflows. Expert corrections. Outcomes someone verified.
The work is turning that production experience into evaluations, training data, and better agents the customer controls. We'll keep publishing what we learn, including the parts that don't work yet.
We welcome feedback, discussion, and research collaborations.
Read all 57 pages, or start with Figure 1.
Read the PreprintSources and further reading
- Recursive Self-Improvement in AI: A Survey, Wu et al., Preprints.org, posted September 30, 2026 (preprint, not yet peer reviewed), doi:10.20944/preprints202609.2680.v1
- χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?, arXiv:2605.16679, May 2026
- Cura 1T: Specialized Model for Agentic Healthcare, arXiv:2607.15314, July 2026
- χ-BENCH Update: Frontier Agents Complete Only 28% of Complex Healthcare Workflows, Actava.ai, June 2026
- The Flywheel Only Spins Where You Can Check the Work, Actava.ai, September 2026
- Meet Weiran Yao · Meet our Advisors: Caiming Xiong
- Actava.ai KORA · Actava.ai Compliance · Actava.ai χ-BENCH · Actava.ai CURA

Written by
Weiran Yao
CAIO & Co-Founder
AI research leader with deep expertise in large-scale model training, reasoning engines, and multi-agent systems. At Salesforce AI Research, Weiran led the post-training team behind the xLAM large action models, along with the Agentforce planning engine and the CodeGenie agent, and drove the pre-training of Salesforce's CoDA code diffusion model. He recently spearheaded open-source releases including Enterprise Deep Research, Webscale-RL, and CoDA-1.7B, advancing applied AI from research into production at enterprise scale. A CMU Ph.D. in Machine Learning, he has published 70+ papers, holds 13 U.S. patents, and serves on NeurIPS, ICML, ICLR, and ACL committees.


