The vast majority of the automated tests for Wolverine, Marten, Polecat and the rest of the Critter Stack almost inevitably touch databases and message brokers and very frequently have to deal with asynchronous behavior. Our tests often have to bootstrap and tear down .NET IHost instances in tests. CritterWatch is even more challenging in that its tests often involve asynchronous messaging between 4-5 different IHost instances. Unsurprisingly, we've struggled mightily with flaky tests both locally and in CI to the point where we didn't even try to run several test suites in CI.

Even after dealing with the obvious issues of data isolation between tests and better conditional "waits" for asynchronous behavior, we still had tests that were unreliable -- especially in CI where the runner images have far less computational horsepower than your development boxes. We never entirely stamped out "flaky" tests, and we started introducing some crude test retries in CI and to run problematic tests in smaller clusters at a time, but even so we still had to mark a lot of tests to be skipped in CI -- which is never an ideal circumstance. I was also becoming very frustrated with CritterWatch development with minimal insight into long running test suites that were sometimes turning out to be hung processes or effectively dead because of unrelated Docker issues.
To finally beat our automated testing woes once and for all, we built (admittedly vibe-coded) the new Bobcat.Supervisor library I want to introduce today.
Dmytro Pryvedeniuk made some absolutely heroic efforts that greatly improved our CI and test automation even before this AI-heavy Bobcat effort
Bobcat is a new JasperFx toolset for writing, supervising, and running integration tests in .NET. This post covers just one piece of it, the supervisor (Bobcat.Supervisor on NuGet). The supervisor takes an existing test suite, potentially splits it across worker processes, and reports what actually happened, including the flaky tests, the crashed workers, and the test that hung.
The supervisor already runs all of Wolverine's CI and CritterWatch's test gate, and most of the numbers below come from those two codebases. Bobcat has allowed us to make our integration test suites far more reliable, faster in some cases, and it has also helped us to identify exactly which tests are flaky in CI for detailed research later.
In our usage, Bobcat.Supervisor enables us to:
- Identify which tests are flaky, which has enabled us to zero in on those particular tests and fix some real problems in core library code in addition to improving test code
- Do selective retries of certain tests depending on why they failed. It's an "allow list" in this case
- Detect when test processes are hung and kill the process -- which has happened to us with CritterWatch from time to time. #sadtrombone
- Selectively restart the entire test process on certain exception types and test failures, which has been a godsend for CritterWatch where tests might spawn over a dozen different .NET
IHostinstances in memory - Selectively restart Docker containers on certain exception types and test failures
- Parallelize some long running test suites by managing multiple parallel tracks of the tests using separate Docker containers for some resources. That's done through environment variables very similarly to how Aspire manages Docker for you
- I won't get into that in this post, but
Bobcat.Supervisoralso works with another forthcoming Critter Stack tool called "Stoat" to give us visibility into progress with very long running test suites so I don't pull out my hair wondering if the AI agent is dead - Analyze some memory usage problems to help diagnose certain test failures. CritterWatch has occasionally had some memory pressure issues in some of our usage, and this functionality has helped solve those issues
Alright, the rest of this is purely AI generated, but here's a rundown of what Bobcat.Supervisor does so far. It's in a pre-1.0 state right now, so there's admittedly room for improvement on usability. The key point really is that we're heavily dogfooding Bobcat's supervisor and it's more than paid off for our own development and for Wolverine as a community project.
Bobcat.Supervisor is also covered in the critterstack-sdd-bobcat-authoring skill of the JasperFx AI Skills pack, as of version 1.14.
Your tests don't need to know about it
The supervisor never loads your test assembly. It launches the compiled test executable and drives it over the wire protocol of Microsoft.Testing.Platform (MTP). So your test project needs no reference to Bobcat at all. It can be plain xUnit v3 in our case, and it works the same on a Bobcat spec project. The only requirement is that the test project is an MTP host.
This is how Wolverine turns it on for every test project in the repository, in its Directory.Build.props:
<PropertyGroup Condition="$(MSBuildProjectName.EndsWith('Tests'))">
<UseMicrosoftTestingPlatformRunner>true</UseMicrosoftTestingPlatformRunner>
</PropertyGroup>That property only changes the executable's entry point. dotnet test, --filter, TRX output, and coverlet all keep working as before. We confirmed that on Wolverine before switching its CI over. Set it in Directory.Build.props rather than on individual CI jobs: a build without the property still produces an executable at the same path, and the supervisor can't drive it.
The supervisor runs from whatever drives your build. For both Wolverine and CritterWatch, that's a Nuke build script. Here's the heart of Wolverine's version, from build/SupervisedTests.cs, trimmed a little:
We'll make this smoother over time!
var factory = new MtpWorkerFactory(executable)
{
EnvironmentFor = laneEnvironment(workers, postgresDatabasePerLane, sqlServerDatabasePerLane),
OnBeforeKill = captureBeforeKill,
BeforeKillTimeout = TimeSpan.FromMinutes(10)
};
var supervisor = new Supervisor(factory)
{
MaxParallelWorkers = workers,
TestFilter = shardFilter is null ? NotFlaky : t => NotFlaky(t) && shardFilter(t),
RetryBudget = DisableTestRetry
? RetryBudget.None
: new RetryBudget { MaxAttemptsPerTest = 3, MaxRetriesPerRun = MaxRetriesPerRun },
ReleaseIdleLanes = true,
Log = message => Log.Information(" {Message}", message),
StallThreshold = TimeSpan.FromMinutes(5),
HeartbeatInterval = TimeSpan.FromSeconds(30),
ResourceSampleInterval = TimeSpan.FromSeconds(15),
KnownTestDurations = knownDurationsFor(projectName, framework)
};
if (!DisableTestRetry) supervisor.AddFailurePolicy(new RetryFailuresInFreshProcess());
var results = supervisor.Run().GetAwaiter().GetResult();The rest of this post goes through what each of those settings does. TestFilter is a plain predicate over each discovered test, including its traits, so Wolverine's standing exclusion of known-flaky tests is just this:
static bool NotFlaky(WorkerTest test)
=> !(test.Traits.TryGetValue("Category", out var category) && category.Contains("Flaky"));--disable-test-retry on the build turns retries off completely. Wolverine measures with retries off first, because a retry budget can hide exactly the instability that parallelism might introduce.
CritterWatch's build started as a copy of Wolverine's, and where it differs is telling. Its test gate had been failing about one run in three on main with no changes at all, so CritterWatch turns retries off by default. The first job was to measure the real flake rate, and retries would have hidden it:
var supervisor = new Supervisor(factory)
{
MaxParallelWorkers = workers,
TestFilter = shardFilter is null ? notChaosRecording : t => notChaosRecording(t) && shardFilter(t),
// Opt in with --enable-test-retry. Zero retries is the honest baseline.
RetryBudget = EnableTestRetry
? new RetryBudget { MaxAttemptsPerTest = 3, MaxRetriesPerRun = 25 }
: RetryBudget.None,
ReleaseIdleLanes = true,
PublishToMonitor = true,
StallThreshold = StallThresholdMinutes > 0 ? TimeSpan.FromMinutes(StallThresholdMinutes) : null,
StallAction = OnStall, // KillAndRetry by default, more on that below
HeartbeatInterval = TimeSpan.FromMinutes(1),
ResourceSampleInterval = TimeSpan.FromSeconds(30),
Log = message => Log.Information(" {Message}", message)
};That notChaosRecording filter comes from a lesson of its own. CritterWatch excludes its chaos tests from the gate because they depend on the chaos monkey actually producing failures. That exclusion is correct, but it used to happen silently. The nightly job that was supposed to run those tests got switched off, and five tests ran nowhere for seven weeks while every record showed a complete-looking 511 tests. Now the filter writes down every test it removes, and the build prints the whole list as a warning:
var chaosExcluded = new SortedSet<string>(StringComparer.Ordinal);
bool notChaosRecording(WorkerTest test)
{
if (NotChaos(test)) return true;
chaosExcluded.Add(test.DisplayName ?? test.Uid ?? "(unnamed)");
return false;
}Split by class, never by test
The supervisor discovers the suite, then deals it out to worker processes by test class. It never splits individual tests across processes, and that rule is there for correctness. Every test framework's isolation contract is per class or per collection, so a class's fixtures and static state assume they're all in one process.
We learned this the hard way on Wolverine. Splitting its persistence tests per test failed one to four tests at random on every run. Splitting per class passed 78 of 78 in the same wall clock. The culprit was a class that built its schema name from a static int counter. Split across four processes, each counter restarts at zero, and all four workers collide on the same schema. Nothing in the test list gives that away, so the partitioner has to be conservative.
Each worker gets its own database
Keeping a class together can't stop two different classes that share a database from landing on different workers. The supervisor gives you a per-worker hook for that. Each worker process gets a Lane number, and you hand each lane its own environment:
new MtpWorkerFactory(executable)
{
EnvironmentFor = worker => new Dictionary<string, string>
{
["WOLVERINE_POSTGRES"] =
$"Host=localhost;Port=5433;Database=wolverine_w{worker.Lane};Username=postgres;password=postgres"
}
}That's essentially the code in Wolverine's build, which creates wolverine_w0 through wolverine_wN up front and points each lane at its own copy. It does the same for SQL Server. The only change the suite itself needed was to read its connection string from an environment variable instead of a hard-coded constant. After that, a CI target just asks for workers and per-lane databases:
Target CISqlServer => _ => _
.ProceedAfterFailure()
.Executes(() =>
{
var sqlServerTests = RootDirectory / "src" / "Persistence" / "SqlServerTests" / "SqlServerTests.csproj";
LaunchDockerServices("sqlserver");
BuildTestProjects(sqlServerTests);
AwaitDockerServices("sqlserver");
RunTestProject(sqlServerTests, workers: 4, sqlServerDatabasePerLane: true);
});The build also caps the worker count at half the machine's cores. A GitHub-hosted runner has 4 vCPUs and 16GB, so those 4 workers become 2 there. That cap exists because the runner was killed outright when too many test hosts ran at once, which brings us to the Marten suite.
A suite that starts its own container per process, like Testcontainers from a [ModuleInitializer], already has this isolation for free. Wolverine's Redis tests went parallel with no source changes at all.
The speedups so far
| Suite | Tests | Sequential | Parallel | Speedup |
|---|---|---|---|---|
CritterWatch SqlServerTests | 205 | 1488s | 314s (6 workers) | 4.7x |
| Polecat.Tests | 1587 | 954s | 366s (4 workers) | 2.6x |
| Wolverine PersistenceTests | 78 | 164s | 73s (4 workers) | 2.2x |
| Wolverine Redis.Tests | 144 | 452s | 206s (4 workers) | 2.2x |
Every one of those parallel runs was fully green, with no retries needed.
Not every suite should be split inside one CI job, though. Wolverine's Marten tests were the 18-minute long pole of its CI matrix. Four worker lanes, each on its own database, cut that to about 9.5 minutes. Then the GitHub runner itself was killed twice in a row on an unrelated change, which is what running out of memory looks like there: several Marten test hosts plus Postgres don't fit on a 16GB runner. So Wolverine moved the parallelism out of the runner and into the CI matrix. The same TestFilter predicate now splits the Marten suite across three CI jobs, each running sequentially against one database:
Target CIMartenDistribution => _ => _
.ProceedAfterFailure()
.Executes(() => runMartenShard(includeNamespaces(MartenDistributionNamespaces)));
Target CIMartenTenancy => _ => _
.ProceedAfterFailure()
.Executes(() => runMartenShard(includeNamespaces(MartenTenancyNamespaces)));
// Everything the other two shards don't claim, so a brand new namespace
// lands here automatically instead of silently dropping out of CI
Target CIMarten => _ => _
.ProceedAfterFailure()
.Executes(() => runMartenShard(
excludeNamespaces([..MartenDistributionNamespaces, ..MartenTenancyNamespaces]),
alsoSubscriptions: true));The split is balanced by measured time per namespace, and the three shards come out within 0.2 seconds of each other. The first attempt balanced by test class count instead, and that turned out to be nearly backwards: the most expensive namespaces were the ones with a few slow tests. And because an empty filter would otherwise pass quietly, Wolverine's build fails any shard whose filter matches zero tests. A renamed namespace can't turn into a green job that ran nothing.
These numbers also flatten out quickly, and the reason is simple. The largest test class sets a floor that no number of workers can get under. The ceiling is simple arithmetic:
ceiling = sum(all test durations) / largest class's total durationWolverine's Redis suite has 599 seconds of test time and one 188-second class, so it can never beat about 3.2x. Eight workers measured 199 seconds, right on the prediction. CritterWatch's SqlServerTests had lots of headroom, but the gains still tapered off: going from 1 to 4 workers saved 1,074 seconds, 4 to 6 saved 100, and 6 to 9 saved only 28 seconds for 50% more test hosts on one SQL Server container. So CritterWatch settled on 6.
The supervisor helps you decide this with data. A run's RunReport ranks the slowest tests by their share of wall clock and reports parallel efficiency. You can also feed the previous run's per-test durations back in through KnownTestDurations, and the supervisor balances lanes by measured time instead of test count. Before Wolverine did that, a count-balanced run had one lane finishing at 101 seconds and another at 11.
Wolverine closes that loop in CI. Every run writes each test's first-attempt duration to a JSON file, the run on main publishes those files as an artifact, and the next run downloads them and passes them back in:
static IReadOnlyDictionary<string, TimeSpan> knownDurationsFor(string projectName, string framework)
{
if (!Directory.Exists(PreviousDurationsDirectory)) return null;
var file = Directory
.EnumerateFiles(PreviousDurationsDirectory, durationsFileName(projectName, framework),
SearchOption.AllDirectories)
.FirstOrDefault();
if (file is null) return null;
var raw = JsonSerializer.Deserialize<Dictionary<string, long>>(File.ReadAllText(file));
if (raw is null || raw.Count == 0) return null;
return raw.ToDictionary(
pair => pair.Key,
pair => TimeSpan.FromMilliseconds(pair.Value),
StringComparer.Ordinal);
}Returning null is fine. With no history, the supervisor balances by count, and any test it doesn't have a duration for is charged the median of the ones it does. It records first-attempt durations because a total that includes retries would overweight exactly the flaky tests.
Retries that don't hide anything
A retry budget is easy to build and easy to abuse. Ours is built so that retried tests can't quietly disappear into the pass count:
A pass on retry is not a clean pass.
SupervisorResultskeepsCleanPasses,PassedOnRetry,Failed, andIndeterminateas separate lists, and the build logs every retry pass as[FLAKY].Retries run in a fresh process. Wolverine tried warm-process retries first and got retries that could never succeed, because a failed first attempt leaves half-built state behind in the process. A failure policy decides what each failure earns:
csharpclass RetryFailuresInFreshProcess : IFailurePolicy { public Disposition Decide(AttemptContext attempt) { if (attempt.Succeeded || !attempt.RetriesAvailable) return null; return Disposition.RetryInFreshProcess( "a failure is retried in a fresh process, within the budget, to separate flaky from broken"); } }A crashed worker never looks like a pass. Tests a crashed worker never reported on come back as indeterminate, listed by name, along with the worker's exit code and last lines of stderr. They're never dropped, and they're never counted as passes or as ordinary failures.
The data this produces paid off right away. Wolverine's Azure Service Bus job was green for four straight runs on main while spending 22 of its 25 retries every time. The same 22 tests failed on their first attempt and passed alone in a fresh process, because the emulator wasn't warmed up yet. That job accounted for 85% of all the flakiness in the repository, and in the GitHub UI it looked exactly like a job with zero retries.
Wolverine now turns every SupervisorResults into a retry ledger entry:
var entry = new LedgerEntry
{
Job = ledgerJobName,
Project = projectName,
Framework = framework,
Tests = results.Tests.Count,
CleanPasses = results.CleanPasses.Count,
PassedOnRetry = results.PassedOnRetry.Count,
RetriesPerformed = results.RetriesPerformed,
Failed = results.Failed.Count,
Indeterminate = results.Indeterminate.Count,
WorkerFaults = results.WorkerFaults.Count,
AbortReason = results.AbortReason,
FlakyTests = results.PassedOnRetry.Select(x => x.DisplayName).Take(MaxRetriesPerRun).ToArray(),
FlakyFailures = results.PassedOnRetry.Take(MaxRetriesPerRun).Select(firstFailure).ToArray(),
IsPartial = results.IsPartial,
Stalled = results.StalledTests.Count,
PeakWorkerRssMb = RunResources.For(results).PeakBytes is { } peak ? peak / (1024 * 1024) : null
};Each entry goes three places: a markdown table in $GITHUB_STEP_SUMMARY, so the numbers are on the run page; a ::warning annotation whenever any retries happened; and a JSON file that a roll-up job compares against the last completed run on main. The change in the retry count is the signal worth watching. FlakyFailures keeps the error from each flaky test's first attempt. Without it, the ledger could name a flaky test but not say why it failed, and two of Wolverine's flaky tests were undiagnosable for exactly that reason.
The ledger deliberately doesn't fail the build past some retry threshold. A suite sitting at three retries today would be one bad day away from a red main. The value is in making the retries visible.
CritterWatch goes one step further. Every supervised run writes a JSON record, and the supervisor's own RunReport.ToJson() is the payload. The build only adds the facts the supervisor can't know, such as which suite this was and whether retries were on:
var record = new JsonObject
{
["suite"] = projectName,
["at"] = DateTimeOffset.UtcNow.ToString("O", CultureInfo.InvariantCulture),
["workers"] = workers,
["retries"] = EnableTestRetry,
["durationMs"] = results.Duration.TotalMilliseconds,
["report"] = JsonNode.Parse(RunReport.ToJson(results))
};./build.sh FlakeLedger then rolls every record up into a committed, generated FLAKE-LEDGER.md: one table per suite and one per test, listing everything that was ever less than a clean pass along with how many runs it appeared in. The current ledger covers 2,431 suite runs from 341 invocations of the build. "This test is flaky" is now a measured rate with a denominator, not an impression. It also keeps runs with retries on and runs with retries off as separate experiments, since a flake shows up as a failure in one and as a retry pass in the other. Writing the record can never fail the run, either. A measuring tool that can break the thing it's measuring is worse than no tool at all.
Naming the test that hung
The worst CI failure is a job that hits its time limit and tells you nothing. Before the supervisor, a wedged Wolverine Marten job's last log line said "275 test(s): 275 batched, 0 isolated", printed 18 and a half minutes before the 20-minute cap cancelled the job. The log couldn't tell anyone which test hung.
Wolverine used to cover this with a shell watchdog that guessed at stalls from outside the process. The supervisor now tracks all of it directly, and every one of these settings only reports by default:
var supervisor = new Supervisor(factory)
{
// Name any test that's been in flight for more than five minutes, with its lane and pid
StallThreshold = TimeSpan.FromMinutes(5),
// One progress line every 30s, led by the longest-running test
HeartbeatInterval = TimeSpan.FromSeconds(30),
// Sample worker memory every 15s and attribute growth to tests
ResourceSampleInterval = TimeSpan.FromSeconds(15),
// Only relaunch lanes that are actually needed, so idle test hosts don't pile up in memory
ReleaseIdleLanes = true
};On top of that:
StallActiondecides what happens after a stall. The default isReport.KillAndRetrykills the wedged worker, gives the stalled test one retry in a fresh process, and sends the other tests from that batch back to their lane.AbortRunends the run. A stall never counts against the retry budget and is never reported as flaky, because a hang is a different problem from a flake. AfterMaxStallKills(3 by default) kills, the next stall aborts the run, on the assumption that the infrastructure itself is broken.MtpWorkerFactory.OnBeforeKillhands you the live worker process, with its pid, right before the supervisor kills it. It's only called for a worker that refused to exit, so a healthy run never pays for it. Wolverine uses it to rundotnet-dumpand print the async stacks, and that capture is what diagnosed a set of wedged Pulsar producers:csharpstatic Task captureBeforeKill(WorkerKillContext context) { if (context.ProcessId is { } pid) { Console.WriteLine( $"[stall] a live worker (pid {pid}) is about to be killed — {context.Reason}. " + "Capturing its async stacks first."); captureAsyncStacks(pid); // dotnet-dump collect, then analyze with dumpasync --coalesce } return Task.CompletedTask; }Supervisor.Snapshot()returns the partial results of a run that was cut short. Wolverine calls it from a SIGTERM handler, so a job cancelled at its time limit still writes its ledger and names its stalled tests before the runner kills it:csharpvar snapshot = supervisor.Snapshot(); // IsPartial = true recordLedger(projectName, framework, snapshot); Log.Warning("=== {Project}: {Summary} ===", projectName, snapshot.Summarize()); foreach (var stalled in snapshot.StalledTests) { Console.WriteLine( $"[stall] stalled at cancellation: {stalled.DisplayName} " + $"({(int)stalled.InFlight.TotalSeconds}s in flight, lane {stalled.Worker.Lane}, " + $"pid {stalled.Worker.ProcessId?.ToString() ?? "unknown"})"); }One catch Wolverine found live: the GitHub runner only signals the step's own shell, and nothing in the
bash→build.sh→dotnet runchain passed SIGTERM along. The build now writes its pid to a file so a small relay script can signal the right process.ISupervisorObserverlets you listen in while the run happens. Its members all default to no-ops, so you only implement what you need. Wolverine's CI prints a line as each lane picks up a test, so even a run that never finishes shows what each lane was doing when it stopped:csharpclass NarrateTestsAsTheyStart : ISupervisorObserver { public void TestUpdated(WorkerLaunchContext worker, WorkerTestUpdate update) { if (update.InProgress) Log.Information(" [lane {Lane}] -> {Test}", worker.Lane, update.DisplayName); } } if (Environment.GetEnvironmentVariable("GITHUB_ACTIONS") == "true") supervisor.AddObserver(new NarrateTestsAsTheyStart());
Wolverine leaves StallAction at Report. CritterWatch switched to KillAndRetry, and the reason is a good argument for having a supervisor at all. A SignalRTests worker wedged for two days. The dump showed a lock inversion inside the .NET runtime itself, between an assembly load and a JIT allocation, after which every thread that needed to JIT anything queued up behind them. CritterWatch's own 90-second host-start timeout and xUnit's ten-minute test timeout were both useless, because no managed timeout can run in a process whose JIT is deadlocked. Only something outside the process could end it, and the supervisor sat healthy the whole time, waiting.
Before turning on the kill, CritterWatch checked the obvious worry, that a five-minute kill would eventually hit a test that was just slow. Its run records held 467 runs with 7 stalls, and every one of those had run to exactly 600 seconds, meaning xUnit's timeout had already killed it. Not a single test had ever gone past five minutes and then passed. So the kill costs nothing that wasn't already lost, and it captures a dump that the xUnit timeout never did.
A killed test that passes on its solo retry counts as a clean pass in the results, which is right for the ledger. CritterWatch still prints each one prominently, so nobody skimming for green misses that something hung:
foreach (var kill in results.StallKills)
{
Log.Warning(
" [STALL-KILLED] {Test} — killed after {Seconds}s in flight on lane {Lane}{Pid}; retried alone. "
+ "Dump (if captured): artifacts/dumps/",
kill.DisplayName,
(int)kill.InFlight.TotalSeconds,
kill.Lane,
kill.ProcessId is { } pid ? $" (pid {pid})" : "");
}Memory sampling came with a lesson of its own. Memory growth during a test measures heap expansion, not live objects, so whichever test happens to be running when the GC grows the heap gets blamed for all of it. Five runs of the same unchanged Wolverine suite named five different "worst" tests, while peak RSS stayed flat. Wolverine logs the top three plus the peak, and treats the peak as the number that matters:
var memory = RunResources.For(results);
if (memory.IsMeasured)
{
var growth = memory.TopRetainers(3)
.Select(x => $"{RunResources.Delta(x.RetainedBytes!.Value)} {x.DisplayName}")
.ToArray();
Log.Information(" peak worker RSS {Peak}; largest RSS growth during: {Growth}",
RunResources.Humanize(memory.PeakBytes!.Value),
growth.Length > 0 ? string.Join(", ", growth) : "(none attributed)");
}A test that really does hold on to memory keeps its place in that top three across runs. If all three names are new every time, the ranking doesn't mean anything.
Finding order-dependent tests before they find you
Some tests only pass because another test ran first. Going parallel exposes those without having caused them, and they're miserable to track down from a red CI run. IsolationSweep looks for them up front. It runs every test class in its own process, compares the results against a normal full-suite run, and sorts whatever changed:
OrderDependent: passed in the full run, failed alone. This is the bug the sweep exists to find.InterferenceVictim: the same kind of defect, seen from the test that got broken.FailedInBoth: an ordinary red test.EnvironmentSensitive: failed during the sweep but passed on a serial confirmation run. That's the sweep's own contention, so it's reported separately and not counted against the suite.
CritterWatch runs a sweep from a Nuke target. The findings are a to-do list, not a pass/fail gate, so the target only fails if the sweep itself couldn't finish:
var sweep = new IsolationSweep(new MtpWorkerFactory(testHostFor(suiteProject(SweepSuite))))
{
MaxParallelWorkers = TestWorkers ?? 4,
Log = message => Log.Information(" {Message}", message)
};
var results = sweep.Run().GetAwaiter().GetResult();
if (results.AbortReason is not null)
throw new InvalidOperationException($"Sweep aborted: {results.AbortReason}");
foreach (var f in results.OrderDependent)
Log.Error(" [ORDER-DEPENDENT] {Test} — {Error}", f.DisplayName, f.ErrorMessage);
foreach (var f in results.InterferenceVictims)
Log.Warning(" [INTERFERENCE-VICTIM] {Test}", f.DisplayName);
foreach (var f in results.FailedInBoth)
Log.Warning(" [FAILED-IN-BOTH] {Test} — {Error}", f.DisplayName, f.ErrorMessage);
foreach (var f in results.EnvironmentSensitive)
Log.Warning(" [ENVIRONMENT-SENSITIVE] {Test} — {Error}", f.DisplayName, f.ErrorMessage);Plan on a sweep taking roughly six times the suite's serial run time at 4 workers.
CritterWatch's rule is that no suite gets more than one worker until its sweep comes back clean. SqlServerTests went to 6 workers only after the sweep reported 205 tests in 52 classes with no order dependence, no interference victims, and no baseline failures.
Watching a run live
A supervised run can also publish itself as it goes: lanes starting and finishing, each worker's pid, retries with the reason for each, stalls, crashed workers, and per-test progress, even for a plain xUnit suite with no Bobcat reference. Set PublishToMonitor = true to turn it on. The publisher is fire-and-forget and switches itself off if nothing is listening, so it can never slow down or fail a test run. That's why CritterWatch's build leaves PublishToMonitor = true on unconditionally: when nothing is listening, the cost is one refused local connection. The whole suite shows up as a single run, not one per worker process, because the supervisor passes BOBCAT_RUN_ID and BOBCAT_RUN_OWNER to every worker. The full event list is in What a Run Publishes.
Getting started
dotnet add package Bobcat.SupervisorThen read these, in this order:
- Reliable Integration Testing: the walkthrough, from "be an MTP host" to per-worker isolation and reporting flakiness honestly.
- Making a test suite safe to run in parallel: the three bugs your first parallel run will probably turn up (a connection string rewritten with a text replace, server-scoped names shared by every worker, and test identities that depend on the environment), plus what to check first when a run goes red.
- Integrating Bobcat with CI: exit codes, filtering, and publishing runs from a pipeline.
All of the Wolverine code above is real, and the complete versions live in its build/ folder: SupervisedTests.cs has the supervisor setup and per-lane databases, CITargets.cs has the per-job targets and the Marten shards, RetryLedger.cs has the ledger, TestDurations.cs has the duration feedback loop, and StallCapture.cs has the dump capture and cancellation handling. The CritterWatch excerpts come from its own Nuke build, which started as a copy of Wolverine's.
A word of warning from both rollouts: the first parallel run of an established suite is a triage exercise, so plan time for it. The supervisor doesn't create the bugs it finds. A test that shares a database name with another test, or that only passes after something else has run, was already broken. Running in parallel just makes that impossible to miss. Get the suite reliable first, then make it fast.
Bobcat is MIT-licensed and on GitHub. Questions and war stories are welcome in the Critter Stack Discord.



