Activation-based monitoring for enterprise AI agents
Reliability monitoring that doesn't burn tokens.
Clearwood reads model intent to catch failures that slip past your monitors and pass review, at a fraction of the cost of an always-on LLM judge.
Already in an active design partnership with a major market-intelligence & research firm.
Our probe-to-judge architecture runs in production at frontier labs like Anthropic and Google DeepMind. We're building the version to catch failures in your business and deployment context, for closed- and open-weight.
Catch failures your current tools don't.
Observability doesn't show you when a model was convincing but wrong. These are examples, not exhaustive.
Confident answers that aren't grounded.
An assistant answers questions over your proprietary documents. Most answers are right. Failures read with the same confidence, but the model has misunderstood.
Data biased toward the answer you want.
A model turns unstructured notes and reports into structured data. It shades ambiguous calls toward the outcome it knows the organization is hoping for. Each one looks defensible. The aggregate is biased.
The change that reads fine, and isn't.
Generated code is idiomatic and locally correct, so it clears review, and the tests pass, because the agent wrote code that satisfies them.
Monitoring that reads intent, not just output.
A near-free probe, trained on the model's internal state, scores every action. Only the small fraction it flags escalates to an LLM judge. You get always-on coverage without paying to run a second oversight model.
Your agent acts
Its transcript and actions stream out as it works.
Probes score every action
Classifiers read internal state and flag suspicious behavior.
Judge only the flagged few
An LLM judge reviews the ~5% that probes surface, not everything.
Attest or intervene
Evidence and alerts, or re-prompt and hold for review.
Early results.
First pass results. The foundation we're building on.
We achieve AUROC of 0.99 when we judge only the top 5% that the probes flag. With 90x more tokens an always-on LLM judge achieves 0.992.
Deception detection signal holds across reader model size with no sharp degradation.
of refit performance retained when probes transfer zero-shot to five unseen but conceptually related agentic tasks.
Three steps to value.
Three steps, each with standalone value. Start small, on your own data, with minimal integration. We work within your data governance constraints.
Retrospective audit
You get a found-failures report and a savings estimate, measured on your own data.
Shadow monitoring
You get precision and recall measured on your production stream, with nothing in the critical path.
Design partnership
You get a monitored deployment, converting to an ongoing engagement on success.
Who you'd work with.
A small team of researchers and operators. Enterprise standards, startup speed.
Matt Levinson
PhD in statistics / computational biology. Scaled Nike's ML organization from 6 to ~100, then ~1.5 years in AI safety and mechanistic interpretability research.
Michael Klear
ML engineering leader shipping production systems at Nike, Coinbase, Yofi (acquired) and Cognitive Space (acquired). Went deep on agents and interpretability before co-founding Clearwood.
Scott Burns
Startup-to-enterprise operator. Co-founded Guide Financial (acquired) and Twine, launched WeaveGrid's first utility products. Builds evaluations and benchmarks (AMLBench).
Start with a retrospective audit.
Share your challenging workflow or monitoring problem. We'll come back in two weeks with findings on failure detection, challenges, and areas for investigation. Clearwood is a public-benefit corporation. Integrity and safety are written into the charter.