
There's a new critter brewing in the Critter Stack. Bobcat is an MIT-licensed tool for authoring, supervising, and running integration tests in .NET, and it has two big jobs in our world:
- To be the Critter Stack's answer for Spec Driven Development, where executable specifications sit in between an Event Model (or just a conversation with an AI agent) and the Wolverine and Marten, Polecat, or Fisher code that implements it -- but let me say that we're treating "Spec Driven Development" mostly as the old idea of "Behavior Driven Development," but now with AI!
- To make the integration testing we do to build the Critter Stack itself more reliable and a whole lot easier to troubleshoot when something does go wrong -- then make that available to everything else!
We will introduce Bobcat in our Event Modeling the Critter Stack Way video later this week on the JasperFx Software YouTube channel if you'd rather watch than read:
Yes, "Bobcat" breaks our Mustelidae naming convention. It's been my working codename for a modernized successor to my old Storyteller tool for longer than the "Critter Stack" name has existed. Where I grew up, a "Bobcat" is also any skid steer, and that carries a nice connotation of just getting stuff done.
I've already written about one piece of Bobcat. The Bobcat Supervisor is the part that drives an existing test suite from outside the process to split it across workers, retry selectively, flush out the flaky tests, and name the test that hung. That one doesn't even require your test projects to reference Bobcat. This post is about the rest of it: how you write specifications with Bobcat, why its output looks the way it does, and where it fits in Spec Driven Development.
Just know upfront that Bobcat is pre-1.0 and we're still moving pieces around between it and Wolverine. I'll call out the parts that are still in flight as I go.
Projecting specifications from the tests you already have
Bobcat was initially going to be "just" another Gherkin tool, but I think I'm personally very disillusioned with Gherkin/FitNesse/Gauge type tools that tried to make it possible for non-technical users to create specifications that could then be turned into automated tests. Instead I think we're shifting the attention to letting you use just plain code that can be "projected" into user readable specifications for collaboration with domain experts.
I think the most interesting idea in Bobcat is that you don't have to write your tests in a special tool to get specifications out of them. If you already have xUnit.net v3 or TUnit tests, Bobcat can project a readable specification from those tests through a little bit of (hopefully judicious) usage of attributes, comments, and source generators. Your tests stay where they are and keep running on their own test runner.
Here's the shape of it, in the style of Marten's async daemon tests:
using Bobcat.Xunit; // or Bobcat.TUnit
[BobcatFeature("Async daemon"), BobcatScenario]
public class when_the_daemon_catches_up : DaemonContext
{
[Fact]
public async Task the_projection_catches_up()
{
// Given the events are published
await PublishEvents();
// the daemon polls on its own schedule, which is why the wait below exists
await WaitForNonStaleAsync();
// When the projection daemon is running
using var daemon = await StartDaemon();
// Then every expected aggregate matches
await CheckAllExpectedAggregatesAgainstActuals();
}
}Three of those comments are steps in the specification and one is just a comment. A "marker comment" is any // comment that starts with Given, When, Then, And, or But, and everything else is left alone. That matters when you're pointing Bobcat at a real test suite that's already full of explanatory comments. If Given/When/Then doesn't fit the test, a // * bullet makes any sentence a step.
The second way in is to decorate the shared helper methods that your tests are already calling:
[Given("the events are published on {threads} threads")]
internal Task PublishMultiThreaded(int threads) => …Now every test that calls PublishMultiThreaded(3) renders "Given the events are published on 3 threads" as a step, with real timing and its own pass or fail. Bobcat does that with C# interceptors generated at compile time, so there's no reflection or stack walking going on at runtime. The two approaches compose too. The comments give a test its narrative, and the decorated helpers report the actual work underneath each sentence.
The constraint that shaped all of this was that a big existing test suite has to be able to opt in one class at a time, with no mandatory base class, no signature changes, and no rewrite. Marten's DaemonTests project was the first real guinea pig. Those tests now render as Bobcat specifications without any of the test bodies being edited.
When you run a projected suite from a terminal, it prints the specifications that it just executed. This is from Bobcat's own sample suite, where some of the specifications fail on purpose:
Feature: Calculator
════════════════════
arithmetic holds OK
✓ Start with the number 5
✓ Multiply by 3 then add 4
✓ The number should now be 17
✓ number: 17
asserting values FAILED
✓ For X=2 and Y=3, the Sum should be 5 and the Product should be 6
✗ For X=4 and Y=4, the Sum should be 6 and the Product should be 8
✗ Sum: expected '6', got '8'
✓ Product: 8
✓ For X=1 and Y=1, the Sum should be 2 and the Product should be 1
2 specification(s) — not greenA few things to know about the projected model:
- It's xUnit.net v3 or TUnit today through the
Bobcat.XunitandBobcat.TUnitadapters. xUnit v2 can't hand a test's result to an attribute, so it's not going to be supported - NUnit and MSTest adapters are on the list, but they're coming later
- A marker comment declares a step, it doesn't execute one. The comment says what the test claims to be doing, and the decorated helpers report what actually ran. Bobcat keeps those two things separate on purpose, and the documentation on projected specifications is pretty blunt about the limits
The Storyteller lineage
Bobcat is the spiritual successor to Storyteller, and to FitNesse before that. Storyteller as a project was admittedly a bust, but there were some ideas in it that really did work. I've said before that my coding superpower is having a longer attention span than average, and Bobcat is the result of something like 17 or 18 years of work and thinking in this space.
The biggest idea worth keeping was making table driven and data centric tests efficient to write and easy to read. So many of the integration tests we write for the event stores or for messaging are really about a lot of data: set up a couple dozen rows of state, do something, then check an entire collection of results. Written as one assertion per value, those tests stop being readable long before they stop being useful.
Bobcat brings over Storyteller's tables as inputs, set verifications, decision tables, and "check this object against this row" grammars. They all work from plain C# tests, where a table is just a raw string literal:
[Then("the unordered details should be")]
internal void TheUnorderedDetailsShouldBe(StepTable expected)
=> SetVerificationComparer.Verify(_details, expected, keyColumns: "Name");_sets.TheUnorderedDetailsShouldBe("""
| Amount | Date | Name |
| 10 | TODAY-2 | Socks |
| 200 | TODAY-1 | The Pants |
| 100 | TODAY | The Shirts |
""");And this is the important part. A failure isn't "expected collection A but got collection B, good luck." It's a grid that tells you which rows were wrong, which were missing, and which were extra:
✗ Then the unordered details should be
╭───┬─────────────────────────┬──────────────────────┬────────────┬─────────╮
│ # │ Amount │ Date │ Name │ Status │
├───┼─────────────────────────┼──────────────────────┼────────────┼─────────┤
│ 1 │ expected '11', got '10' │ 2026-09-29 (TODAY-2) │ Socks │ FAIL │
│ 2 │ 200 │ expected …, got … │ The Pants │ FAIL │
│ 3 │ 100 │ TODAY │ Sweatpants │ MISSING │
│ 4 │ 100 │ 2026-10-01 │ The Shirts │ EXTRA │
╰───┴─────────────────────────┴──────────────────────┴────────────┴─────────╯Rows are matched up by their key columns first and then compared, so a wrong value in a matched row is one failed cell in that row. A markdown table pastes in unchanged, so the table in the GitHub issue, the table in the pull request, and the table in the test can be the exact same text. There's much more in the Data Intensive Specifications tutorial.
Test diagnostics an AI agent can use
There has to be an AI hook of course, and in this case it's purposely building in AI-friendly diagnostics and telemetry to hopefully make it much easier for your AI agents to solve test failures along the way
That same lineage leads straight to what I think might be the most valuable thing about Bobcat right now. The other Storyteller idea that worked was having the test harness itself write out what the system did during a test instead of only whether the final assertion passed. Way back when, that meant embedding the message history of a service bus right into Storyteller's HTML results so you could understand a failure without attaching a debugger.
Fast forward to today, and it's very often an AI agent reading your test failures first. When all the agent gets from a failing integration test is Expected True but was False, its next move is to add some logging and run the test again. That's an extra round trip that a better failure message would have saved, and with tests that run against real databases and message brokers, those round trips are slow.
This is the worst with asynchronous tests, which is of course exactly what we're doing all day with Wolverine and the event stores. Think about a typical Wolverine test. You send a message, that message cascades other messages, some events get appended, an async projection eventually catches up, and then you check the outcome. When that test fails, the evidence you need doesn't belong to any one assertion. It's the entire causal chain of what was sent, what was received, what was retried, and what ended up in the dead letter queue. Wolverine's tracked sessions already know that whole story, and a test that only asserts on the outcome throws all of it away.
Bobcat's answer comes in a few parts:
- Every disagreement, not just the first one. A normal assertion library throws on the first failure and you never find out about the others. Bobcat's
SpecAssertgathers the results, and the test still fails once at the end. You can also opt into having your existing Shouldly assertions projected as steps, where a run of consecutive assertions is all evaluated before the next action - Failures as data. Expected and actual values are reported in cells and grids like the one above, not as one long exception message that somebody has to parse
- Scenario reports. A specification can attach a table describing what the system did, and by default it's only written out when the scenario fails
That last one is the direct descendant of Storyteller's custom reporting. Here's a trimmed version of the worked example from Bobcat's own test suite, shaped on Wolverine's envelope records:
public class MessageActivityReport : TableReport
{
public override string Title => "Message activity";
// One envelope record, stating what happened and judging nothing
public void Record(long atMs, string envelopeEvent, string message, string? destination, int attempt)
=> Row(("at (ms)", atMs), ("event", envelopeEvent), ("message", message),
("destination", destination), ("attempt", attempt));
}var report = SpecReport.For<MessageActivityReport>();
report.Record(0, "Sent", "ConfirmAppointment", "local://confirm", 1);
report.Record(3, "Received", "ConfirmAppointment", "local://confirm", 1);
report.Record(19, "MessageSucceeded", "ConfirmAppointment", null, 1);That same code works from a Gherkin specification or a projected xUnit.net or TUnit test. For a projected test, the report lands in that test's own output from dotnet test, which is the first place a developer or an agent is going to look when a test goes red. There's also a JSON report from the command line runner for a more machine friendly format.
Reports are capped at 200 rows with a visible "and N more" line. That's a lesson from experience, as a real test suite once dumped a 698MB tracked session log into its test output and took down the whole test worker with it. Context isn't free, and a test that dumps everything is just as unreadable as a test that dumps nothing.
The Wolverine specific helpers for this, including the tracked session report, are moving out of Bobcat and into a new WolverineFx.Bobcat library that will live in the Wolverine repository so that Wolverine dogfoods them in its own tests. That's still on the way, so today the Wolverine integration is the Bobcat.Wolverine package.
The Gherkin model
Sometimes Given/When/Then really is the right way to express a specification. Maybe you've got non-developers who want to read the specifications, or folks on your team who are already comfortable with that approach. So Bobcat also supports Gherkin, with one big difference from SpecFlow or Reqnroll. Bobcat compiles your .feature files. A Roslyn source generator turns every step into a direct method call at build time, so there's no runtime reflection and no step registry. Better yet, a step that doesn't match anything is a compiler error instead of a pending test that you find out about later.
For Critter Stack applications, Bobcat ships with a grammar for event sourcing and messaging. This specification is from the BankAccountES sample in the Bobcat repository:
@domain:Banking
Feature: Freeze Account
@slice:FreezeAccount
Scenario: Freezing an account records the freeze
Given no events for Account "77777777-7777-7777-7777-777777777777"
And events for Account
| Event | AccountId | ClientId | Currency |
| AccountOpened | 77777777-7777-7777-7777-777777777777 | 88888888-8888-8888-8888-888888888888 | USD |
When FreezeAccount is received
| AccountId | Reason |
| 77777777-7777-7777-7777-777777777777 | Suspected fraud |
Then AccountFrozen is emitted
And the Account read model contains
| IsFrozen |
| true |And here's all of the code behind it:
public class FreezeAccountFixture : WolverineCritterStackFixture;Everything else comes from the shipped grammar. Account, FreezeAccount, and AccountFrozen are resolved to your real .NET types at compile time. The When step sends the command through a Wolverine tracked session, so the Then steps only run after every cascading message has finished. The arrange and assert steps go through the JasperFx.Events abstractions, which means the same specification runs against Marten, Polecat, or Fisher without any changes. The grammar today covers these steps:
| Step | What it does |
|---|---|
Given events for {aggregate} / Given {event} occurred | Arranges the history on an event stream |
When {command} is received | Sends the command and waits for everything it caused |
Then {event} is emitted / Then no events are emitted | Checks what the command appended |
Then the {readmodel} read model contains | Checks a projected view, one verdict per column |
Then {message} is sent | Checks the outgoing messages |
Then the command is refused / Then validation fails with {string} | Checks the unhappy paths |
Of course you can write your own grammars for whatever your system needs. There's a document database flavored module in the box too, plus Alba support for specifications that go through HTTP endpoints.
A couple of other things I like about this model:
A Bobcat specification project is a Microsoft.Testing.Platform host, so every scenario shows up in
dotnet test, your IDE's test explorer, and CI just like any other testdotnet run -- previewshows you every step and exactly which method it's bound to without executing anything. It doesn't even need your database to be running:○ Given the left operand is 25 ↳ CalculatorFixture.TheLeftOperandIs — "the left operand is {int}" value ← "25" (capture)The same
[Given],[When], and[Then]attributes work for both models. A step method that's matched against a.featurefile today can be called directly from an xUnit.net test tomorrow and it renders the same way
The Behavior Driven Development with Gherkin tutorial walks through this from a brand new project.
Bobcat's role in Spec Driven Development
Now to the main reason that Bobcat exists at all. I laid out the bigger Event Modeling and Spec Driven Development strategy for the Critter Stack a couple weeks back. The short version is that there are a lot of ways to arrive at a design for a system. You might have a formal Event Model, an export from the EventModelers.AI platform, some low fidelity C# stub types, or you might have just talked it through with an LLM. However you got there, something has to pin down what "done" means for each vertical slice before the code gets written, and it needs to be something that both people and AI agents can read and execute. That's Bobcat's job.
The workflow we're building toward looks like this:
Start with a model. Commands, events, aggregates, and views can begin life as empty stub types, with the Event Model declared in code. Bobcat's
bobcat import-event-modelcommand can write those stubs for you from an EventModelers.AI board exportWrite the specifications against the stubs. They're red until the real code exists, which is the whole point. A specification that isn't even ready to run can be marked as pending, and that shows up on the Event Model as an open question instead of a failing test:
csharp[BobcatSpec(typeof(ConfirmAppointment), Pending = true)] public async Task a_confirmed_appointment_cannot_be_confirmed_twice() => throw new NotImplementedException(); // never runsImplement until the specifications go green. By hand, or more likely now with an AI agent using our AI Skills, which include skills for scaffolding specifications from a model and for going from a failing specification to the implementation in idiomatic Wolverine and Marten code
Keep the model and the code honest. Every specification has one stable identity, and it can be bound to the slice of the Event Model that it's evidence for. That's the
@slice:FreezeAccounttag in the Gherkin up above, or[BobcatSpec(typeof(ConfirmAppointment))]in a projected xUnit.net test
That last step is what makes me think this could really work. As the real code grows, the Critter Stack derives the actual Event Model from your handlers, endpoints, and projections. Anywhere the model you declared disagrees with the code you wrote, or a slice has nothing specifying it at all, that shows up as a hotspot on the model. Bobcat even has an audit you can run to catch a specification that nobody designed and a design that nothing specifies.
It's also why the diagnostics I talked about above matter so much. Spec Driven Development with AI agents is only as good as the feedback loop. An agent can only work its way from a red specification to a green one efficiently if the red specification tells it what actually happened. A specification that says "an AccountFrozen event was expected, and here are the events that were actually appended and the messages that were actually sent" is something that an agent can act on in one pass.
I'd be lying if I said that every piece of this was finished today. The scaffolding of code skeletons is moving out of Bobcat and into a Wolverine command, the declared Event Model is moving into a fluent interface in JasperFx.Events, and Stoat is going to be the user interface for previewing and running Bobcat specifications alongside the Event Model visualization. I'm also still a little bit skeptical about Event Modeling as a formal "specify everything, then generate the system" method. I'm old enough to remember Model Driven Development! But executable specifications that keep a model honest against real code are something I'm very comfortable betting on, and that's the part that Bobcat owns.
Making Critter Stack development more reliable
Even if Spec Driven Development isn't your thing, Bobcat has already paid for itself in the more mundane work of building the Critter Stack. Almost all of our tests touch databases and message brokers, and most of them have to deal with asynchronous behavior. As I wrote about in the Supervisor post, we've struggled mightily with flaky tests and slow builds, and I've pulled out plenty of hair wondering if a long running test suite was hung or just slow.
Here's where we're using Bobcat in our own work today:
- The Supervisor runs all of Wolverine's CI and CritterWatch's test gate, and it's in the build for Marten and Polecat as well
- Marten's async daemon tests are being rendered as projected Bobcat specifications. Those are some of the most complicated asynchronous tests we have, so they were where better step by step diagnostics would help the most
- The CritterCrush sample application in our CritterStackSamples repository is the proving ground for the whole Event Model to specifications to code workflow
- The next step is the
WolverineFx.Bobcatlibrary I mentioned above, so that Wolverine's own messaging and event sourcing tests get the same Given/When/Then helpers and tracked session reporting that we want you to have
I realize that some of this might turn out to be mostly useful for developing the Critter Stack itself. I'm okay with that. More reliable tests for Marten and Wolverine are a direct benefit to everybody using them, and I'd much rather hand you integration testing tools that we've been beating on ourselves every single day.
Trying it out
Bobcat is on GitHub under the MIT license, and the packages are on NuGet. For the projected model with xUnit.net v3:
dotnet add package Bobcat
dotnet add package Bobcat.Generators
dotnet add package Bobcat.XunitThe documentation is at bobcat.jasperfx.net, and I'd start with whichever of these sounds like your week:
- Specifications with Code for projecting specifications from xUnit.net or TUnit tests
- Behavior Driven Development with Gherkin for the
.featurefile model - Data Intensive Specifications for tables, set verifications, and decision tables
- Reliable Integration Testing for the Supervisor
Again, Bobcat is pre-1.0 and we're knowingly making breaking changes as we find out what real applications need from it and trying to use it in our own work. If you do kick the tires, come tell us what worked and what didn't in the Critter Stack Discord. And if you'd like help with Spec Driven Development or your own integration testing, JasperFx Software is always happy to talk.


