Last month the Rails Foundation published something the Ruby on Rails community had not seen before: a leaderboard of coding agents scored on real Rails work. The first Agents on Rails report topped out at 92% in the Accuracy column for the best model tested, which is roughly what you would expect from a field of frontier models. The Rails API recall column tells a different story: it shows the share of runs where the model reached for the Rails API that a task was built around, instead of hand-rolling its own version. In that first report it ran from 8% to 35%, and a follow-up published on September 2 moved the top of the range to 41%.
Those results describe Writebook , the application the benchmark runs against. They do not describe your application. On August 24 the Rails team open sourced lemans , the harness behind every number they have published, which means you no longer have to take the leaderboard’s word for anything.
In this article, you will learn what lemans measures, how to build a small bench you can run on your own machine, and how to prove that bench is worth trusting before you spend anything on it.