SANITY·CHECK
for macOS
Multi-model deliberation

Chatbots don’t earn confidence. They’re built with it.

Right or wrong, they sound exactly the same. The goal isn’t a faster answer. It’s finding where you’re wrong while being wrong is still cheap.

See how it works ↓
You ask: “How many r’s are in strawberry?”
MODEL A · BLIND
There are 2 r’s in strawberry.
MODEL B · BLIND
There are 3 r’s in strawberry.
MODEL C · BLIND
There are 3 r’s in strawberry.
!
Two answers match. One doesn’t. That’s a lead, not a verdict.
Sanity Check surfaces the disagreement. You decide what to challenge and what to verify.
How it works

Put a panel in the room. Keep yourself in the chair.

Sanity Check gives two to eight AI advisors the same question, preserves their independence, and lets you control when the debate begins.

01
Ask once
Every advisor answers cold. Nobody sees anyone else’s response, so the blind round captures genuinely independent takes instead of a chain of peer influence.
02
Compare, not average
Answers land side by side, raw and inspectable. Agreement is evidence to examine—not truth by vote—and disagreement shows you where to spend your attention.
03
Decide with the room
You chair the debate: name the disagreement, press the sharp answer, add missing context, or send advisors back for another round. Peer influence begins only when you say so.
Built for productive friction

The disagreement is the product.

INDEPENDENCE

The blind round is sacred.

Pick advisors across vendors. Three models from one lab can still be one opinion in a trench coat. Diverse blind answers give you a better chance of seeing the weak claim.

THE PROVOCATEUR

A shit stirrer with a job description.

It takes no position. It reads the panel and fires three questions at the premise nobody checked, the constraint nobody named, or the answers that only look like agreement.

NO BLACK BOX

Every answer stays a first-class turn.

Inspect the raw response and the byte-exact prompt each advisor received. Nothing is quietly merged, smoothed over, or hidden behind an aggregate answer.

HUMAN IN THE LOOP

You decide what survives.

Sanity Check organizes the evidence and exposes the fault lines. It never promotes the majority to “truth,” and it never makes the final call for you.

The Decision Brief

A clean ending, written by someone who wasn’t in the argument.

When the dust settles, an independent moderator—never one of the panelists—turns the documented session into a brief you can act on. The Provocateur’s questions, and what they changed, are preserved verbatim.

  • Recommendation
  • Consensus
  • Disagreements
  • 1–10 alignment score
  • What would change the answer
  • Next step
Private by design

Your questions. Your context. Your disk.

PRIVACY PASS

Secrets leave as aliases—or not at all.

Names, contacts, and sensitive details are swapped for aliases on your Mac before a prompt goes anywhere. You see the exact outgoing payload and approve it first.

LOCAL TRANSCRIPTS

Nothing phones home.

Sessions stay on your disk. No Sanity Check account, telemetry, or subscription. Export the full reasoning trail as one Markdown file you own.

Use the models you want

The check comes from the disagreement, not the horsepower.

Cheap and free models can expose what one frontier model says with a straight face. Mix vendors when you can; independence matters more than stacking brand names.

OpenRouter

One key reaches hundreds of models at pay-per-token rates. Build a mixed-vendor panel, watch the meter, and run a real session for a few cents.

Ollama

Run the whole panel locally for no API cost. Keep model traffic on your Mac and choose the local models that fit the job.

Sanity Check. Where disagreement is a feature, not a bug.