What happened
Vals AI, an independent evaluation lab, released its Recursive Self-Improvement Index on Tuesday, September 9. The benchmark, per CryptoBriefing's report, tests whether a frontier model can carry out the full loop of designing a successor architecture, generating training data, running the training run, and evaluating the resulting model against its own baseline. That is a stack of capabilities that until recently sat firmly in the theoretical column of AI safety literature.
Vals AI positioned the index as a diagnostic, not a leaderboard, though scores for individual labs were included. The firm did not publish full methodology on launch day, saying a technical paper will follow. What it did publish is a scoring rubric across four sub-tasks and a summary chart showing measurable, non-zero results for the top-tier models tested.
Why it matters
Recursive self-improvement, or RSI, has been the load-bearing assumption behind a decade of AI-risk arguments. If a model can build a stronger model without a human in the loop, the pace of capability gains stops being gated by hiring cycles at frontier labs. It becomes gated by compute.
That is a different world for regulators, and a different world for the crypto-adjacent AI infrastructure trade that has priced compute scarcity as a durable tailwind. The RSI Index is the first attempt to put a number on where we actually are on that curve. The number is not zero.
It is also not one. That middle ground is exactly the range where governance gets hard, because the models are capable enough to matter and not capable enough to obviously demand a pause.
