NeoSavant Research · Model Observatory

There is no universally best model. There is a qualified intelligence architecture for a governed job.

METIS evaluates intelligence architectures against the same governed meaning, evidence, and controls — showing not simply which model performs best, but which architecture is qualified for the job.

A NeoSavant Research initiative · Governed Intelligence

What a benchmark misses

Same interaction. A plausible answer — and the governed one.

A model can look strong on generic benchmarks and still misread the business. Governed evaluation is where that shows up.

Governed work-task Principal interaction reasonseparate the reason for contact from salient sentiment and side-requests

Synthetic interaction

Customer“I've been on hold twenty-five minutes, third time I've called this week, the service has honestly been unacceptable. Anyway — looking at my statement, I've been charged the $89 activation fee twice and I need one taken off. Also I'm moving in May, so I'll need to update the address at some point.”

Before — without governance

Service complaint — quality

Plausible — and wrong

benchmark-strong model · confidence 0.94 · single label

Keyword salience wins: long hold, “third time,” “unacceptable.” The model classifies the loudest signal and collapses the call into one intent.

After — with governance

Billing dispute — duplicate charge

Correct — the reason for contact

principal reason · context and side-request held separately

The resolvable, business-relevant reason is the double $89 charge. The complaint is sentiment; the move is a future, non-principal request — neither is the reason for contact.

Evidence: “charged the $89 activation fee twice… need one taken off.”

Why it matters: Routed as a “service complaint,” the duplicate charge is never corrected and the ticket closes unresolved. The governed reference separates reason, sentiment, and future request — the distinction downstream analytics and resolution actually depend on.

NeoSavant synthetic exercise. Model labels illustrate the failure mode, not a specific product's output.

METIS IQ · a governed measure of intelligence performance

How well a system does a defined governed job — with the evidence beneath the score.

Model-neutral. Each architecture is scored on the dimensions meaningful to its workload; what doesn’t apply is marked, not invented. The readout below demonstrates what the assessment produces; qualified model scores publish in the role-fit listing as frozen test evidence completes.

Synthetic example

METIS IQ

87 / 100

Governed contact-center workload · v1

DevelopingQualified band

Qualified for governed reasoning on this workload. The dashed band marks the qualification threshold, not a rank against other models.

Claude — Opus / Sonnet class

Adjudicator & reviewer

API · eval 2026-08-28 · benchmark v1 · dev single run

Governed correctness
9.1
Evidence grounding
8.8
Calibrated refusal
8.5
Governability
9.2
Boundary consistency
8.6
Latency parity
N/A
Evidence coverage82% of workload casesInspect evidence

N/A is a real result: latency wasn’t measured comparably in this run, so it isn’t scored. The score never borrows credit from dimensions that weren’t evaluated.

Role-fit, not a leaderboard

No single best model — a qualified role for each, on the same governed job.

The governed workload needs a team: something to hold the reference, something to run at volume, something to check the checker. METIS IQ qualifies each for the role it’s actually fit for.

Adjudicator & governed reviewer

ClaudeOpus / Sonnet class · API

Pending

METIS IQ

Holds the reference interpretation and settles boundary cases. Strongest on calibrated refusal — knows when the evidence won't support a stronger claim.

First-pass semantic

Qwen · Self-hosted

Pending

METIS IQ

Carries the high-volume first pass and hands ambiguous interactions up to the adjudicator. Governable and cost-efficient at scale.

Cross-family reviewer

Mistral · Self-hosted

Pending

METIS IQ

An independent second read from a different model family — catches correlated errors a same-family reviewer would miss.

Executor

NVIDIA Nemotron · Self-hosted

under-measured

METIS IQ

Throughput executor for settled definitions. Capacity-limited in these runs, so it isn't qualified yet — shown as pending, not padded.

Narrow high-volume signals

Classical ML · Self-hosted

workload-specific

Not a reasoner and not scored as one — reasoning dimensions are N/A. Precise and cheap on bounded, well-defined signals where it belongs.

Reviewer candidate

GPT / OpenAI · API

varies

Bake-off in progress. Ratings publish when the governed run completes — no score is shown until it's earned on this workload.

The listing mirrors the hand-approved model grid. METIS IQ values publish per model as qualification completes — Pending means not yet qualified, never an estimate.