Skip to content
Svi članci Gems

The agents write the code. Your team decides what gets built.

30. rujna 2026. · 15 min čitanja · Ivan Blažević

plan_driven is a Ruby gem that runs a Rails team's delivery process from the terminal. An interview becomes an implementation plan grounded in your real code. You approve it, it becomes tickets, and each ticket goes to its own AI agent, which opens a pull request. Nothing merges until you've read it and every acceptance criterion has a passing test. We used it to build one feature in a new Rails app, and recorded all of it.

Download this article as a PDF

Coding agents got good this year. Give one a clear task and a codebase and it will write the migration, the model, the controller, the views and the specs, and open a pull request before your coffee is cold. At Rubycode we use them every day.

What didn't get better is everything around the code. Which feature, exactly? What happens when the event is full, or the organizer tries to RSVP to their own event? How do you split it so that each pull request can be reviewed in one sitting? Who agreed to the plan, and where is it written down? When the agent says it's done, how do you know every requirement has a test, and not just the easy ones?

Those questions used to be answered by a process: a plan in Confluence, a review, tickets in Jira, pull requests linked to tickets, and QA ticking off acceptance criteria. Agents made the code fast, so the process is now the slow part, and on most teams it quietly disappears. The agent gets a one-line prompt, and the decisions a team would have made together get made by the model, in a tool call nobody reads.

We didn't want to give up the speed or the process, so we wrote the process down as code.

A plan first, then one agent per ticket

plan_driven adds a command, plan-driven, to your Rails app. It takes a feature through the same phases a careful team already uses, and it won't let a phase start before the one before it is approved:

  1. Interview. You answer the questions only people can answer: what, why, where, who, when, background and what's out of scope.
  2. Plan. A model drafts the rest of the implementation plan from your answers, your schema and your code: the existing data structure, architectural and database changes, risks, security, monitoring and testing. You read it as a PDF, change what's wrong and approve it.
  3. Tickets. The approved plan is split into tickets with acceptance criteria, estimates and dependencies. You steer them, approve them, and they become GitHub issues.
  4. Development. Each ready ticket goes to its own Cursor cloud agent, which opens a pull request. A ticket starts only when the tickets it depends on are merged.
  5. Review. You read each pull request, run the review checks, send feedback to the same agent if needed, and approve and merge from the terminal.
  6. Evidence. Every acceptance criterion is mapped to the Cucumber scenario that proves it, and a delivery report records what was built, by whom, and who approved it.

The model does the writing. The rules are Ruby. That split is the whole idea, so we'll come back to it.

Or click through it in the browser

New in 0.2.0 is a browser wizard, mounted at /plan_driven in development. It has five steps: Plan, Approve, Tickets, Agents & PRs, and Proof. It doesn't reimplement anything. Every button runs one plan-driven command, and a panel beside the form shows that command and its output as it runs. The guards are the same, the audit trail is the same, and anything you click you could also type.

We delivered a second feature in Gather with it, comments on events: six tickets, six agents, six pull requests, one round of feedback, and 42 of 42 acceptance criteria proven by a passing scenario.

The wizard, 12 minutes, narrated: the interview as a form, redrafting one section of the plan, the tickets, six pull requests, feedback to one agent, the evidence and the feature in the app. After every click, the terminal panel shows the command that ran.

One feature, from idea to production code

To show it honestly, we built a new Rails 8 app for it, Gather. It's a small events site for the Zagreb tech community: you can browse events, and organizers can create them. Events have a capacity, but nothing counts seats, so organizers collect names in chat and the small rooms overflow.

The feature was RSVPs with a waitlist. Signed-in people RSVP and cancel, the event page shows seats left, a full event puts you on a waitlist, a cancellation promotes the first person waiting, and the organizer sees who's going. It's small enough to follow in one sitting and has enough edge cases to be interesting: two people racing for the last seat, the organizer's own event, what "cancel" does to the waitlist.

The planner and all five agents ran on Claude Opus 5.5, through a Cursor account. Here is the whole run, recorded as it happened. Nothing is staged: the plan, the tickets, the pull requests and the report are all in the Gather repository.

The CLI deep dive, the whole run in a terminal, 28 minutes, narrated: the interview, reading the plan, changing a section of it, the tickets, five pull requests read and merged, two rounds of feedback, the evidence and the finished feature.

An interview, not a prompt

plan-driven new "RSVPs with a waitlist" asks seven questions. I typed nine lines in total, and none of them mentioned tables, models or controllers. Those are the model's job.

The draft came back with eight assumptions listed under the plan: cancelling deletes the RSVP rather than marking it cancelled, the waitlist is first come, first served, the organizer sees names but not email addresses, and so on. That list is the first thing to read, because it's where the model made a decision on your behalf.

The interview in the terminal, and the assumptions the model made
Figure 1. Seven questions, then a draft. Every assumption the model made is listed, so you know what to check first.

Reading the plan

The plan is written to docs/plans/pd-1-rsvps-with-a-waitlist/ as Markdown, HTML and PDF. It reads like one an engineer on your team would write, because it is grounded in the code: it cites Event#organized_by?, the require_organizer filter in the events controller, and the fact that production runs on SQLite, so a race for the last seat needs an IMMEDIATE transaction. Before I saw it, a guard had already checked that every model, file and column in Existing Data Structure really exists.

The implementation plan, open in the browser
Figure 2. The plan follows the implementation plan template many teams keep in Confluence, and cites only models, files and columns that exist.

The model had a list of outstanding questions, and they were the right ones. Can the organizer RSVP to their own event? When does RSVP close? Should the events index show seats left too? Those are product decisions, and the model rightly didn't make them. I answered them in one sentence each, and plan-driven redraft rewrote that section with a Decided list, down to which method enforces each decision.

The decisions recorded in the plan
Figure 3. My answers, recorded under Decided. The questions that don't block the work stay open.

Outstanding questions is only one section. Every section of the plan can be changed: What, Why, Database changes, Application changes, Risks, Testing. The draft is the model's proposal, and the team has the final say. To show it, I drafted a second plan for Gather, PD-2, event categories. The model proposed a category string column on events. We'd rather have a table, so plan-driven redraft PD-2 database_changes "Use a categories table instead of a string column…" rewrote that one section: a new categories table, seeded with four rows, and a category_id reference on events, still in expand and contract steps. For a small change I didn't need the model at all. plan-driven edit PD-2 database_changes opened the section in vim, and I added a color column by hand.

Adding a column to the plan in vim
Figure 4. edit opens any section in your own editor. Saving makes a new revision, and the guards run again.

Every change is a new revision in the log, the plan goes back to draft, and the guards check it the same way they check the model's draft. Approvals belong to a revision, so a changed plan is approved again. Once tickets are drafted, the plan is locked and changes go into a follow-up plan.

Then submit, which runs the guards again, and approve. Revision 3 was approved. If anyone edits the plan now, it needs approving again.

Tickets you can steer

The first breakdown had nine tickets and one of them had ten acceptance criteria. A guard flagged it. The model had split the work by user action, while I wanted it split by layer so that each pull request builds on the last one. plan-driven tickets PD-1 "Make it five tickets…" redrafted them: the migration, the model rules, the RSVP card, the organizer's list, and seats left on the index.

The five tickets in the plan's work overview
Figure 5. The tickets are added to the plan's PDF, with their dependencies, estimates and acceptance criteria, so they're read where the plan is read.

approve-tickets created a GitHub issue for each one, with the story, the acceptance criteria and the implementation notes straight from the plan. That issue is what the agent works from, and what I'll check its pull request against.

A ticket as a GitHub issue
Figure 6. Each issue carries the plan key and the ticket, so anyone can trace it back to the approved plan.

One agent per ticket, in order

plan-driven develop PD-1 started one agent, for T1, because the other tickets depend on it. Each agent gets the ticket, the parts of the plan it needs, our team's conventions, and the rules its pull request will be checked against: which Cucumber file to write, how to tag each scenario, what it may and may not change. plan-driven prompt PD-1/T1 shows the exact text, so there is no hidden prompt.

Reading the pull request

When it finished, the agent opened pull request #8. I read it on GitHub, like any other pull request: the migration with its foreign keys, the unique index and the status check constraint, the regenerated schema, and a Cucumber feature with one scenario per acceptance criterion.

The agent's pull request on GitHub
Figure 7. The migration ticket's pull request: the schema change, and a scenario for each acceptance criterion.

Then plan-driven review PD-1/T1 checked what a reviewer shouldn't have to check by hand.

plan-driven review, every check passing
Figure 8. The review checks: the ticket and issue, the size, the specs, that a migration ticket only changes the database and tests, a scenario for every criterion, and green CI.

approve-pr records my approval and posts it on the pull request. merge asks me to type the ticket key, merges, and says what's ready next. For the pull requests with UI, I checked out the branch and clicked through it before approving, because the checks prove the criteria have tests, and only a person decides whether the feature is right.

Trying the RSVP card on the branch
Figure 9. The RSVP card on the T3 branch, tried locally before it was approved.

Feedback goes back to the same agent

Two pull requests needed another round. T5, seats left on the index, was opened before T3 was merged, so its CI had run without T3's code. The review caught that the branch was two commits behind main. The helper also repeated a rule that already lived in Event#seats_left. One plan-driven feedback command sent both points to the same agent, which pushed to the same pull request, and the next review was clean.

A behind-main warning, and feedback sent to the agent
Figure 10. The review notices the branch is behind main. The feedback asks the agent to merge main and to move a duplicated rule onto the model.

Proof, not promises

With all five merged, plan-driven evidence PD-1 ran the plan's Cucumber scenarios on main. Each of the 33 acceptance criteria had a passing scenario, including the one where several threads race for the last seat and exactly one wins.

33 of 33 acceptance criteria proven
Figure 11. Every acceptance criterion, and the scenario that proves it.

plan-driven report PD-1 wrote the delivery report: each ticket with its pull request, merge commit and approver, the proof matrix, the guard findings, every approval and the full timeline. It's the document I'd send to a product owner, and nobody wrote it by hand.

The delivery report
Figure 12. The delivery report, generated from the audit trail.

The feature

And the feature works the way the plan says. Three people take the three seats, the fourth joins the waitlist as number one, and when someone cancels she's promoted automatically. The organizer sees who's going, in order, with names only.

Maja on the waitlist
Figure 13. The event is full, so Maja joins the waitlist.

Maja promoted after a cancellation
Figure 14. Petra cancels, and the seat goes to the first person waiting.

The organizer's list of who is going and who is waiting
Figure 15. The organizer's view: who's going, in the order they RSVPed, and who's waiting.

On main, Gather now has 71 specs and 49 Cucumber scenarios, all green, and RuboCop is clean.

The rules live in code, not in the prompt

The easiest way to build a tool like this is a long prompt: "always write tests, keep pull requests small, never drop a column in one step". Models follow instructions like these most of the time. "Most of the time" is fine for a draft and not fine for a merge.

So in plan_driven, the model writes and Ruby checks. Guards are plain Ruby classes that run at every step:

  • PlanGuard checks that every required section is written, that security ends with a risk level, and that Existing Data Structure cites only models, files and columns that really exist.
  • MigrationGuard refuses removing or renaming a column in one step unless the plan stages it, and warns about NOT NULL without a default and non-concurrent indexes on PostgreSQL.
  • TicketNormalizer corrects anything with one right answer, like title prefixes, estimates rounded to your scale and dependencies renumbered, and reports each correction.
  • TicketGuard checks acceptance criteria, estimates, dependency cycles, the order of expand and contract on each table, and that every table the plan changes is covered by a ticket.
  • PrGuard checks the pull request's scope and size, that it has specs, that only migration tickets add migrations, that every acceptance criterion has a tagged scenario, that CI is green and that the branch isn't behind main.

When a guard fails on a draft, its errors go back to the model as a list to fix. A plan that still fails isn't accepted. Nothing merges while a guard fails or CI is still running.

Guards don't replace reading the code. They make sure that when you read a pull request, you're reading for design and behaviour, not checking whether someone remembered the tests.

What it asks of you

plan_driven gives your attention to decisions, and takes it away from bookkeeping. In this run, my part was answering the interview, reading the plan, deciding the open questions, steering the tickets, reading five pull requests, trying the UI and sending feedback twice. The agents wrote the code, and the gem kept the paperwork.

It also has requirements, and it's better to know them upfront:

  • The app needs to be on GitHub with CI on pull requests, and the agents are Cursor cloud agents, so you need a Cursor account with GitHub connected.
  • Evidence needs Cucumber. The first agent can set it up if your app doesn't have it.
  • The plan's state lives in the database of whoever drives it. The documents go to docs/plans/, which you commit, so the whole team reads the same plan and report.
  • For drafting, use a strong model. Planning is where it pays for itself. We used Claude Opus 5.5 through Cursor, and OpenAI and Anthropic work too.

Try it

# Gemfile
gem "plan_driven", group: :development
bundle install
bin/rails generate plan_driven:install
bin/rails db:migrate
bundle exec plan-driven configure     # keys go to ~/.plan_driven/config, never into the app
bundle exec plan-driven doctor
bundle exec plan-driven new "Your next feature"

The README walks through setup and the whole run step by step, with the real output. The Gather repository has the plan, the tickets and the delivery report, and pull requests #8 to #12.

It's the first release. If a guard your team relies on is missing, or the drafter got your plan wrong, open an issue. We'd like to hear about it.


About Rubycode

plan_driven is built and maintained by Rubycode, a Ruby on Rails
company from Zagreb. We build and rescue Rails applications, and help teams put AI agents to work
safely. Need Ruby or Ruby on Rails engineers? Get in touch.

Trebate inženjere koji ovako pišu kod?

Povezujemo provjerene programere, dizajnere, QA i voditelje projekata s tvrtkama i startupima diljem EU. Recite nam što trebate i predložit ćemo kandidate koji odgovaraju.

Dogovorite 30-minutni poziv