Run the Rails Agent Benchmark Yourself

Run the Rails Agent Benchmark Yourself

Last month the Rails Foundation published something the Ruby on Rails community had not seen before: a leaderboard of coding agents scored on real Rails work. The first Agents on Rails report opens a new window topped out at 92% in the Accuracy column for the best model tested, which is roughly what you would expect from a field of frontier models. The Rails API recall opens a new window column tells a different story: it shows the share of runs where the model reached for the Rails API that a task was built around, instead of hand-rolling its own version. In that first report it ran from 8% to 35%, and a follow-up published on September 2 opens a new window moved the top of the range to 41%.

Those results describe Writebook opens a new window , the application the benchmark runs against. They do not describe your application. On August 24 the Rails team open sourced lemans opens a new window , the harness behind every number they have published, which means you no longer have to take the leaderboard’s word for anything.

In this article, you will learn what lemans measures, how to build a small bench you can run on your own machine, and how to prove that bench is worth trusting before you spend anything on it.

What the benchmark measures

Two numbers come out of a run, and they measure different things.

Accuracy is the share of runs that pass verification. A run passes when the application’s own test suite stays green and a set of hidden checks pass. Those checks test behavior rather than implementation, so a hand-rolled fix passes exactly like an idiomatic one. Across the models in the first report, accuracy landed between 65.1% and 92.1%.

Rails API recall is scored separately, and it asks a narrower question: did the run reach for the Rails API the task was built around, or did it write its own version of that API? Every task in the corpus, the benchmark’s full set of tasks, turns on exactly one API that the instructions never name. In the first report, recall ran from 8% on DeepSeek to 35% on Claude Fable 5. By the September 2 follow-up, Claude Fable 5.1 had reached 41%, the highest anyone has posted so far.

That spread is a lot wider than the accuracy spread, and it is the more useful of the two if you maintain an application that is a few versions behind. Accuracy tells you whether a model can make a test pass. Recall tells you whether it knows the framework you are actually running. A model that solves the task by writing its own version of something Rails already ships is a model that adds code to your application that you now own, which is the same shape as the hand-rolled helpers you find scattered through most long-lived Rails codebases.

The report also notes that runs that recalled the API solved 92% of the time, compared with 87% for hand-rolled solutions. That gap is smaller than it looks: five points across 21 tasks at three runs per model is a correlation, not a demonstration that reaching for the framework causes correctness, and the models with the best recall tend to be the stronger models to begin with.

A bench you can run locally

lemans is a Ruby gem and it wants Ruby 3.4 or newer:

gem install lemans

The official corpus runs Writebook inside a Daytona opens a new window sandbox, which means an account and a bill. For learning the tool that is more setup than you need, because lemans also supports Docker. Either way a bench is just a directory: a bench.yml holding the run profile, a Dockerfile describing the sandbox, and one directory per task.

The sample bench for this post opens a new window has two tasks against a minimal Rails 8 application:

lemans-sample/
├── bench.yml                    # image, agent, limits, verifier
├── environment/Dockerfile       # a small Rails app with a User model
└── tasks/
    ├── hello-world/
    └── ar-normalize-email/

Each task directory holds an instruction.md (YAML frontmatter plus the request, written the way you would actually file it), a verification_test.rb with the hidden checks that grade the result, and a solution.patch holding a reference fix. A task can also carry an environment.patch to seed its defect, which is how one base image serves many tasks.

lemans tasks reads the bench and lists what is in it:

$ lemans tasks
task                difficulty  tags                             description
ar-normalize-email  easy        rails,active-record,validations  Email lookups should survive the way people actually type email.
hello-world         easy                                         Prove the base image and the grading pipeline end to end

hello-world asks the agent to write Hello, world into a file. It is deliberately not a Rails task. It exists so that when a real task fails, you know the image, the sandbox, the agent install, and the grader were all fine.

Writing a task that measures recall

ar-normalize-email is the one that does the work. Its instruction never names an API:

Support keeps reopening the same ticket. Someone signs up as
`Ada@Example.COM `, comes back a week later, types `ada@example.com`, and we
tell them no such account exists. Worse, they sign up again and now there are
two rows for one person, so the uniqueness check we thought we had is not
holding.

The application ships a User model with nothing but a presence and uniqueness validation on email, so there are two obvious ways to fix this. One is a callback:

before_save { self.email = email.strip.downcase }

The other is normalizes opens a new window , which Rails has shipped since 7.1:

normalizes :email, with: ->(email) { email.strip.downcase }

Both store a lowercased address, so a check that only looks at the column afterwards grades them the same. The hidden checks look at behavior instead. Three of the five checks tell the callback and normalizes apart. Here is one of them:

test "a user is found by an email typed the way a human types it" do
  assert_equal users(:ada), User.find_by(email: " ADA@Example.com ")
end

The callback normalizes on write and leaves the query alone, so this looks for a string nobody stored and finds nothing. normalizes applies the normalization when the attribute is assigned and to the matching arguments of find_by and where, so it covers the write, the finder, and the uniqueness validation together.

Swapping the callback in as the reference solution and grading it makes the difference concrete. It passes the two checks that inspect the stored column and fails the other three, the finder, the where clause, and the duplicate rejection:

VerifierTest#test_a_user_is_found_by_an_email_typed_the_way_a_human_types_it
-#<User id: 225478506, email: [FILTERED], name: "Ada", ...>
+nil

5 runs, 5 assertions, 3 failures, 0 errors, 0 skips

The callback fixes less than half of the ticket and closes it anyway.

A more careful hand-roll wins some of that back. Moving the normalization to before_validation and passing uniqueness: { case_sensitive: false } restores the duplicate check, because the validation now runs against the normalized value. The read path is the part that stays broken either way, and it is the part nobody notices until support reopens the ticket.

That divergence is the part worth copying when you write tasks against your own application. A task where the idiomatic and the hand-rolled fix behave identically measures style, and grading style is a losing game. A task where they come apart measures whether the model knows the framework you are running, which is the thing you actually wanted to find out.

Finding those spots in an unfamiliar codebase is its own exercise, and it is much the same reading you would do to grok a Rails app for the first time opens a new window . The cheapest first task is usually a bug you already fixed: the regression test you wrote becomes verification_test.rb, and the fix you shipped becomes solution.patch.

Proving the bench before you spend anything

A bench can be wrong in two directions, and both are expensive to discover later. If your solution.patch or your checks are broken, every model you benchmark afterwards gets graded against your mistake. If your checks pass on the unfixed application, the task measures nothing at all.

lemans ships two agents that cost nothing and need no API key. The oracle applies solution.patch, and the nop agent does nothing:

lemans run --agent oracle    # every task must score 1.0
lemans run --agent nop       # every task must score 0

Run both before you run a model. On the sample bench, each trial takes about five seconds and costs nothing:

$ lemans run --agent oracle
task                agent   reward  outcome    cost_usd  duration
ar-normalize-email  oracle  1       completed  0         5.3
hello-world         oracle  1       completed  0         5.3

$ lemans run --agent nop
task                agent   reward  outcome    cost_usd  duration
ar-normalize-email  nop     0       completed  0         5.5
hello-world         nop     0       completed  0         5.5

The harness is also careful about the difference between a model failing and your setup failing. A trial that cannot build or start its sandbox is recorded as invalid with an environment_error outcome instead of scoring zero, so infrastructure trouble never quietly depresses a model’s numbers. The same care shows up in verifier.restore, which lists the paths rolled back from a pre-agent snapshot before grading (test, bin, and config/environments/test.rb in the sample), and in preverify, which runs the application’s own suite first. An agent that deletes an inconvenient test does not pass.

Once both sanity agents behave, a real run needs a provider key:

export OPENAI_API_KEY=...
lemans run --task ar-normalize-email --model openai/gpt-5.4 --attempts 3

Three attempts per task is what the official methodology uses, and it is also where the cost comes from. lemans report -A model-task aggregates the results, but on a bench this small the score is the less interesting output. Every trial writes a runs/<model>/<task>__<id>/ directory containing agent.patch, the agent’s entire diff. The score says whether it passed. The patch says how, which on this particular task is the whole question.

I ran GPT-5.4 against it three times. It scored 0 every time, and the three patches show the same instinct: before_validation :normalize_email, a private method doing email.to_s.strip.downcase, uniqueness: { case_sensitive: false }, and a migration that backfills the column and adds a unique index on lower(trim(email)). None of the three reached for normalizes. Two went further and added their own self.find_by_email, which is the part worth sitting with: the model understood that lookups needed normalizing too, and answered it with a finder the rest of the application never calls. The hidden checks call User.find_by(email:), the same thing a controller calls, so all three passed the checks that inspect the stored column and failed the finder and the where clause. Ten to fourteen steps each, and $0.42 for the three attempts.

Conclusion

A bench is three things: a bench.yml, a Dockerfile, and a directory of tasks. Run the oracle and nop agents first, and you know it grades correctly before you spend anything on a model.

Keep what you conclude from it small. The official Stage 1 corpus is made of atomic tasks that each turn on one Rails API, and a two-task sample like this one tells you even less. Neither shows whether an agent can carry a multi-step change through a real application. The Rails team measured that separately in Stage 2 opens a new window : twenty feature tickets against Fizzy, a production kanban app, where the best model solved 35% of its runs.

What a bench like this does tell you is whether a model reaches for the Rails APIs your version actually ships or writes its own, and the further behind your application is, the worse that answer tends to get.

Is your application far enough behind that your agents are writing Rails you retired years ago? We can help. opens a new window

Get the book