Articles on Artificial Intelligence

Run the Rails Agent Benchmark Yourself

Last month the Rails Foundation published something the Ruby on Rails community had not seen before: a leaderboard of coding agents scored on real Rails work. The first Agents on Rails report opens a new window topped out at 92% in the Accuracy column for the best model tested, which is roughly what you would expect from a field of frontier models. The Rails API recall opens a new window column tells a different story: it shows the share of runs where the model reached for the Rails API that a task was built around, instead of hand-rolling its own version. In that first report it ran from 8% to 35%, and a follow-up published on September 2 opens a new window moved the top of the range to 41%.

Those results describe Writebook opens a new window , the application the benchmark runs against. They do not describe your application. On August 24 the Rails team open sourced lemans opens a new window , the harness behind every number they have published, which means you no longer have to take the leaderboard’s word for anything.

In this article, you will learn what lemans measures, how to build a small bench you can run on your own machine, and how to prove that bench is worth trusting before you spend anything on it.

Read more of Run the Rails Agent Benchmark Yourself opens a new window

What Happens to Your Ruby Tools With AI?

When we talk about building with AI, most of the attention goes to what’s new. Models, agent frameworks, protocols, and tools seem to appear every week, making it easy to assume that adopting AI means introducing an entirely new technology stack.

But Rails developers already have many of the building blocks needed to create useful AI-assisted workflows. Generators, Rake tasks, command-line interfaces, schemas, tests, and APIs were designed to make software easier to work with by providing structure and predictable behavior. Those same qualities make them well suited for AI coding tools.

In this context, an AI-assisted workflow doesn’t mean building an agent into a Rails application. It can be as simple as a developer using an AI coding tool to complete a task in an existing codebase, whether that’s adding a feature, running tests, analyzing technical debt, or helping with a Rails upgrade. As these tools become capable of taking more actions on a developer’s behalf, they need reliable ways to interact with the codebase and the tooling around it.

This article explores how existing Ruby and Rails tooling can become part of AI-assisted development, and why introducing AI doesn’t require starting from scratch.

Read more of What Happens to Your Ruby Tools With AI? opens a new window

Running a Ruby MCP Server in Production

In a previous post, AI Assistant for Our Blog Writing Process opens a new window , I introduced the assistant we built to help with our blog writing. At the core of that assistant is an MCP server, which serves as the source of truth for both of our blogs. It exposes that knowledge through tools the client can call and documentation the client can read.

Getting an MCP server running is the easy part. Every quickstart, in every language, gives you a server that runs as a subprocess on your own machine and disappears when the client exits. That’s enough to experiment locally, but it’s a long way from something a team can rely on. Once you want to deploy it, questions about where it runs, state management, authentication, and security become your responsibility. The Ruby SDK’s defaults don’t solve most of those problems, and one of them even comes with a published security advisory.

In this article, we’ll cover what changes when a Ruby MCP server stops being a subprocess: the two shapes it can take in a Rails shop, why session state breaks down when running behind multiple Puma workers, the DNS rebinding vulnerability the transport shipped with, and what the specification asks of you once a shared token is no longer enough.

Read more of Running a Ruby MCP Server in Production opens a new window

Tracking LLM Latency & Cost with Rails Events

Wiring an LLM into a Rails app takes a handful of lines. Understanding what it actually costs you (feature by feature, user by user) is harder. Most providers and SDKs already report tokens, latency, and even cost, but that data lives in their dashboard. It’s disconnected from your requests, your users, and the feature that made the call. And it sits apart from the APM and logs where you already watch the rest of your app.

In a previous post, we introduced Rails.event.notify(...) opens a new window , the tool-agnostic Event Reporter shipping in Rails 8.1. In this post, we’ll put it to work on a real problem: instrumenting every LLM call in your app so token usage, latency, and cost become structured events you can log, graph, and forward to any APM or data warehouse.

Read more of Tracking LLM Latency & Cost with Rails Events opens a new window

From AI Opportunity to AI Feature in Rails

At OmbuLabs.ai, we’ve explored the importance of identifying meaningful AI opportunities opens a new window before selecting a solution. Once a worthwhile opportunity has been identified, however, a new question emerges:

Is this problem worth solving in the first place?

Too often, teams focus on the technology before evaluating the value. AI can automate tasks, generate content, and process information at incredible speed, but if the underlying work doesn’t matter, making it faster won’t create meaningful business outcomes.

Once a worthwhile opportunity has been identified, however, a new question emerges:

What should we build first?

Read more of From AI Opportunity to AI Feature in Rails opens a new window
Get the book