Catching Flaky Tests Before Release: A Sentinel QA Walkthrough

Catching Flaky Tests Before Release: A Sentinel QA Walkthrough
The most expensive test in your suite isn't the slow one. It's the one that fails one run in ten. The build goes red, someone re-runs it, it goes green, and nobody writes anything down. Do that enough times and the team learns the wrong lesson: red doesn't mean broken, it means "try again." The day a red build is real, the first response is a re-run.
This is a walkthrough of one way to stop that habit before it reaches a release: tracking every test's pass rate across runs, and taking the unreliable ones out of the gate while a human decides what to do with them.
The problem isn't the flake, it's the ambiguity
A single failing run tells you almost nothing. It could be a real regression. It could be a timing race, a slow staging environment, a third-party script that loaded late. From inside one run, those look identical.
What separates them is history. A test that has passed forty times and just failed once is very likely telling you something. A test that has passed six times out of ten this week is telling you something different: that the test, not the product, is the thing that's broken. Most CI setups throw that distinction away because they only ever look at the current run.
What Sentinel QA keeps track of
Sentinel QA is an open agentic test runner. On a code change it reads the diff, decides what needs testing, writes the tests, runs them with Playwright, and reports back with Markdown and JSON output and a pass/fail exit code.
The part that matters here is quieter. Sentinel QA keeps a rolling window of pass rates for each test across runs, and uses it to sort every test into one of four states: new, stable, quarantined, or rejected.
- New tests haven't earned a track record yet.
- Stable tests have passed consistently across the window.
- Quarantined tests have a pass rate that says they can't be trusted as a gate.
- Rejected tests are ones that shouldn't stay in the suite at all.
That's the whole idea. It isn't a clever flake predictor. It's bookkeeping, done every run, that most teams mean to do and never get to.
A walkthrough: one week of runs
Say you're shipping a web app and pointing Sentinel QA at each change, something like npx sentinel-qa run --app your-app --diff HEAD~1.
On Monday it generates tests for a checkout tweak. They're new, so they run and report, and nothing is known about them yet. By Wednesday, after a handful of runs, most have settled into stable. One hasn't: a test around a promo-code banner passes on some runs and fails on others, for no change in the code under test.
Here is where a normal pipeline costs you. The banner test fails on Thursday's release candidate, someone hits re-run, it passes, and the release goes out. Nobody learns anything.
With the pass-rate window, the banner test has already drifted into quarantine by then. It still runs, and its result is still recorded, but it no longer decides whether Thursday's build is red. The report shows what was quarantined, so the decision is visible rather than silent. That's the difference between ignoring a flaky test and managing one.
What you do next
Quarantine doesn't fix anything. It moves the problem from "everyone re-runs the build" to "one person looks at one test." Open the report, find the quarantined test, and ask the usual questions: is it waiting on something that isn't ready, does it depend on data another test changes, is the environment simply slow?
Two outcomes are common. If the test was wrong, you fix it and it earns its way back as its pass rate recovers across later runs. If it was never testing anything useful, it gets rejected, and your suite is smaller and more honest for it.
What it does not do
It's worth being exact, because this is a topic where overselling is easy.
Sentinel QA does not have a dedicated flaky-test detector that diagnoses why a test flakes. Quarantine is based on pass rate across runs, not root-cause analysis, so it tells you which test is unreliable and leaves the why to you.
It also doesn't yet plug into your pull requests. There are no PR comments, no GitHub Actions integration and no Slack reporting today; you read the Markdown and JSON output and use the exit code. Testing is web-focused through Playwright, with Flutter support planned rather than shipped.
And because it generates tests from your diff, the run needs an Anthropic API key and has a token budget you set: a run aborts if it goes over max_tokens_per_run, so cost doesn't drift.
Where it fits
If your release process has a quiet rule that "you can ignore that one, it's flaky," this is worth an afternoon. You won't get a flake-free suite out of it. You'll get a suite where unreliable tests are named, counted and kept off the release gate, and where a red build starts meaning something again.
Sentinel QA is an open agentic test runner — code, issues, and roadmap on GitHub: https://github.com/ahn283/sentinel-qa
Frequently Asked Questions
What problem does Sentinel QA address in test suites?
Sentinel QA addresses the issue of flaky tests that intermittently fail, causing ambiguity in build results. It tracks each test's pass rate over multiple runs to distinguish unreliable tests from genuine failures, preventing teams from ignoring red builds due to flaky tests.
How does Sentinel QA classify tests based on their pass rates?
Sentinel QA sorts tests into four states: new (no track record yet), stable (consistently passing), quarantined (unreliable with low pass rates), and rejected (tests that should be removed). This classification helps manage flaky tests by preventing unreliable ones from blocking releases.
What actions should be taken when a test is quarantined by Sentinel QA?
When a test is quarantined, a human should investigate its cause by checking dependencies, environment issues, or test correctness. If the test is fixed, it can return to stable status as its pass rate improves; if it’s not useful, it should be rejected to keep the test suite reliable and concise.
Does Sentinel QA automatically diagnose why a test is flaky?
No, Sentinel QA does not perform root-cause analysis or diagnose why a test flakes. It only tracks pass rates across runs to identify unreliable tests, leaving the investigation and resolution of flakiness to the development or QA team.
What integrations and technologies does Sentinel QA currently support?
Sentinel QA runs tests using Playwright for web applications and plans to support Flutter in the future. It does not currently integrate with pull request tools, GitHub Actions, or Slack; users access results via Markdown and JSON outputs and an exit code. It also requires an Anthropic API key for generating tests from code diffs.
Continue reading

Automating Invoice-to-Spreadsheet Workflows With Pilot on macOS
Invoices arrive as email attachments and leave as spreadsheet rows, and the part in the middle is all you. Here's how that month-end hour gets handed to an agent on your Mac.

Meet Cubist: Learn the Solve, Time It to WCA Standards, Then Compete
Cubist puts the guided path, a competition-grade timer with real inspection rules, and global leaderboards in one 3x3 app — for cubers past their first face.

Meet Nesty: Baby Tracking Built for the 3am Shift
Nesty logs a feed, a sleep or a diaper in two taps — one-handed, in the dark, offline. Wake windows, WHO growth curves, and one shared family log.