An independent, hands-on evaluation by JevResearch · model pin: service-reported jev-1.13.0 · measurements September–October 2026 · report revision 5 (2026-10-03)

A short counterpoint to Archer Hume’s Jev’s Architecture Unmasked.

Jev: Not Frontier, But Still Worth Your Attention

TypeSafe AI sells Jev as a frontier-class reasoner that cannot hallucinate, built by the co-inventor of ChatGPT - fast, and almost free. We ran it live on 16,379 benchmark requests, measured its latency and billing, and probed what it is underneath. The result is a smaller, humbler model that is nonetheless genuinely useful for a job that nobody else serves quite this way.

MMLU-Pro
82.7%
12,032 graduate-level multiple-choice questions
best published: Claude Fable 5.1 92.4% (5-shot CoT)
ARC-Challenge
97.9%
1,172 grade-school science items - a saturated field
best published: Llama 3.1 405B 96.9% (25-shot CoT)
GPQA Diamond
76.5%
196 graduate-level science questions
best published: GPT-6 Astra 96.3% (reasoning enabled)
Humanity's Last Exam (MC)
21.9%
494 expert-exam multiple-choice items
best published: GLM-5.3 62.5% (with tools)
MATH-500 as multiple choice
83.1%
261 encodable items, answer as an option
best published: GPT-5 (high) 99.4% (reasoning enabled)
MATH-500 digit read-out
13.6%
273 items, every digit right or wrong
no protocol-matched reference
ARC-AGI-2, per cell
59.6%
70,100 abstract-puzzle cells - a diagnostic encoding
no protocol-matched reference
ARC-AGI-2, exact grid
0.0%
120 tasks, all cells of all 167 grids correct
best published: GPT-6 Astra 95.0% (reasoning, pass@2, semi-private set)
1 / 8
TL;DR. Jev is not a frontier model; for example, it misses about a quarter of graduate science questions, and scores a mere 21.9% on HLE. It nonetheless is also not a toy: 82.7% MMLU-Pro, 76.5% GPQA, ~73 ms of proxy-reported upstream service time per question, and a full graduate-scale benchmark run for cents. The honest category is cheap real-time sub-frontier judgement - routing, rubric grading, control-plane decisions - where nothing else on the market combines this latency, this price, and a strict output contract with high reliability (19 contract-invalid responses in ~16k benchmark requests).

The pitch, and the parts that survive contact

“A model that cannot hallucinate, at frontier-level performance, built by the co-inventor of ChatGPT - incredibly fast, incredibly cheap, with free output.”

Each clause is technically defensible in a narrow sense and misleading in the sense a buyer will hear. Take them one at a time.

Cannot hallucinate. What Jev actually cannot do is return free text: it is a fixed-output API that answers by selecting among caller-supplied options, so there is no prose to hallucinate in. But a contract-valid answer is not a true answer. Jev is frequently wrong (see our benchmarks), and every one of those wrong answers arrives beautifully formatted. A model that cannot write prose has not solved hallucination; it has made hallucination hard to notice.

Fast and cheap. Both true, with the leading explanation being prefill-dominated scoring without autoregressive decoding.1 Speed and price are properties of that shape of task, not of frontier economics.

Frontier-level. This one simply does not survive. Jev is good, and “good” will turn out to mean something genuinely useful here - but it is not a frontier model by any late-2026 standard, and where it looks frontier-like (97.9% on ARC-Challenge), that's only on a race that finished years ago. Everything below is the evidence.

The co-inventor of ChatGPT. Diogo Almeida certainly deserves credit, but saying “I co-invented ChatGPT” overassigns it. Almeida was one of the eight primary authors on InstructGPT, worked on learned optimizers, and was also one of 88 people thanked in the initial release of ChatGPT. But ChatGPT was born from the work of thousands of people; see the footnote.2

What we did, and what we could not do

We ran Jev live against public benchmarks through its API: 12,032 MMLU-Pro questions (82.7% accuracy), 1,172 ARC-Challenge, 196 GPQA Diamond, 494 Humanity's Last Exam multiple-choice items, the encodable MATH-500 subsets (261 as option MCQ, 273 under a per-digit rubric), and the ARC-AGI-2 public set (167 grids, re-encoded per-cell and as 120 whole-task requests) - 16,379 planned requests in three frozen suites. This repository publishes the derived aggregates and the complete Talk-to-Jev traces. The benchmark item text itself is third-party licensed content and is not republished here. The analysis results are generated from these artifacts.3

Three constraints shape everything below. First, Jev answers directly: one shot, no chain of thought, no tools, no retries - which is not how the frontier scores we compare against are produced with modern models, although it is how older models in the benchmarks functioned, and is automatically factored into the Pareto frontier comparisons. Second, the API is closed: we never saw weights, gradients, or anything internal to its design, so the architecture discussion is inference from observable behavior.3 Third, we chose benchmarks the model could answer natively - selecting between visible options or grading fixed rubrics - because converting them into something else would measure our conversion, not the model.

What Jev appears to be

What follows is our operating premise: every other section reads Jev's behavior through the model stated here. The full ledger - alternative hypotheses with probabilities and a falsifier per load-bearing claim - lives in ARCHITECTURE-ANALYSIS.md.

Leading theory, in one sentence: Jev looks like a small transformer language model, post-trained for judgement rather than conversation, served with its generation head replaced by a probability read-out over caller-supplied options - every question in a request scored from one batched prefill pass, which is why it is fast, why its output is free, and why it cannot write you a poem. The API alone does not identify this uniquely: one pass versus an unexposed equivalent pipeline, a replaced head versus other read-out machinery, a decoder transformer versus other transformer variants, and trained-from-scratch weights versus adapted or distilled ones are all still on the table.
Hypothesized ArchitectureA transformer stack, drawn plain - with the two modifications the evidence actually showsREQUESTState + questionsOptions: 255 or fewer eachChoice / score / noulPOST /v1/systemoneSERVING PATHWhitespace normalizerFixed template, ~316 tokVendor's own tokenizerLatin-centric BPE,byte-level fallback,~1 tok per non-Latin charONE FORWARD PASS (prefill only)Token embeddingsAttentionFeed-forward× NFinal hidden states hEvery question + option reads from this one passCompute: 73 ms floor + 6.1 ms per 1k tokensLinear to 29k; decide cost below resolutionHeavily batched (options × parallel runs)Est. active size: ~4 to 9B paramsREAD-OUT HEADReplaces the language headOne distribution per question,over the caller's options+0.44 ms / question+0.10 ms / optionEach pays only for its own tokensDISPLAY PIPELINEArgmax decided before roundingProbabilities rounded tothe 0.01 grid; returned sumsare 0.99 or 1.00, never aboveRESPONSE JSONSerialized vectorsWHAT MADE THE WEIGHTS (inferred - plausible)Pretraining corpus not identified; knowledge horizonstrongest on 2024-era facts, weak on 2025 newsJudgement-format post-training (vendor: RLCD);Base model ancestry not identified;Frontier-teacher contribution: none identifiable, not excludedNOT IDENTIFIEDDense vs MoE; attention typeTeacher-distilled vs trainedon its own dataExact rule behind the 0.01 rounding
ComponentBest guessConfidence
Output sideAn exposed probability distribution over the caller's options (measured: choice / score / noul as its three exposed shapes).Confident (exposed distribution)
ServingLikely one shared scoring pass: questions and options are scored within a single request, without returned free text. The flat timing under concurrency is measured, but batching versus spare capacity versus replicas is not distinguishedPlausible
Service-time profileFixed ~73 ms intercept + ~6 ms per 1k input tokens on the upstream-service header, approximately linear to 29k tokens, with no large quadratic signature. Measured increments: +0.44 ms per extra question and +0.10 ms per extra option - both consistent with the added tokens' cost.Confident (measured)
Input accountingWhitespace normalization (ASCII runs collapse; NBSP, ZWJ and BOM do not) and a fixed ~316-token template overhead are visible in the reported counts.Confident / plausible
TokenizerThe API's reported token counter follows an English/Latin-centric BPE-style profile: heavy Latin merges, ~1 token per codepoint for the major non-Latin scripts, and a byte-level fallback for uncovered characters (0.99-1.01 tokens per UTF-8 byte on the rarest blocks). The reported counts match no known preexisting tokenizer signature (173 tested); the model's own encoder is not identified by that.Plausible (strong)
CoreTransformer-family; dense vs MoE unknown - with quantized-serving assumptions the two size angles overlap, so neither is forced and MoE stays possible; decoder versus other transformer shapes is not identified at these context lengths, and the attention variant is unknowablePlausible / open
SizeThe service-time slope, read through hardware-dependent scenarios (assumed hardware, quantization, utilization), spans ~0.6 to 49B active parameters. Capability positioning suggests 4 to 14B dense-equivalent. A quantized dense ~4 to 9B is the parsimonious joint reading; a MoE (~15 to 100B total) stays possiblePlausible
TrainingPretraining corpus not identified; on the sampled dated-fact probes the knowledge horizon reads strongest on 2024-era facts and unreliable on 2025 news; judgement-format assistant post-training; OpenAI-flavored brand prior inherited from training text; frontier-teacher contribution: none identifiable, not excludedConfident / Plausible
What the evidence weighs againstNot frontier (measured). Against retrieval/cache assistance and a thin wrapper around another vendor's API. The base model is not identified. See below.Confident (directionally)

See ARCHITECTURE-ANALYSIS.md §1 for details.

What the milliseconds say

In our September timing campaign, responses reported an upstream-service timing header (x-envoy-upstream-service-time), the proxy-reported upstream service time. Our October follow-up responses no longer supplied that header, so those tests use client elapsed time instead. The September sequential probes across prompt sizes from ~0.5k to ~29k tokens fit:

~73 ms fixed floor + ~6 ms per 1k input tokens measured on that header, with a negligible quadratic term and the fit quality reported in the figure. A flat base plus a linear input-token term is what a transformer's forward pass looks like from outside.
0k5k10k15k20k25k30k050100150200250input tokens (thousands)
Proxy-reported upstream service time vs input tokens across 35 probe configurations: a flat floor plus a straight line, with no attention-blowup curvature at these lengths; the fit's residuals and uncertainty stay inside the band drawn on the figure.
serving floor: 73.1 ms (95.6%)prefill (560 input tokens): 3.4 ms (4.4%)
An illustrative decomposition of one representative call under the single-pass fit: most of a typical short question's header time is the fixed floor, and the decision step is small.

The more telling experiment is self-batching. Packing 192 questions into one request barely moves the reported service time (+0.44 ms per extra question), while the reported output grows about 33 tokens per question - each one's full probability vector, serialized. Those marginals are consistent with the input tokens each question (~55) and each option (~18) adds to the shared prompt. Reading those tokens at the prefill slope is itself worth +0.34 ms - the residual 0.1 ms is effectively noise. Scoring 255 options instead of 2 costs +0.10 ms of service time and 9.6 reported output tokens per option, against the +0.11 ms from its own ~18 tokens. In both cases the measured increments are consistent with pure input-token scaling. This is what prefill-dominated scoring without autoregressive decoding would look like from outside, and that is the leading explanation we run with.

The battery re-confirmed this live: across a grid of option counts (2-255) crossed with option lengths (2-64 tokens), nothing scaled with the number of options beyond the tokens they add (-0.8 ms per 100 options, inside the ±20 ms residual noise). Billed output grows ~33 tokens per question and ~9.6 per option - that is the response JSON serializing itself, which is why output can be priced at zero under the service's billing policy. Upstream service time stays flat (78 ms at concurrency 1, 82 ms at 32) while our own wall time bends, consistent with efficient shared serving. A four-question batch answers in ~261 ms where the same four questions separately take ~1,117 ms: no per-question round trip appears on the critical path in these measurements. “System one”, under this reading, is as much a serving description as a psychological one - prefill, read out, done.

0501001502002500326496128160192Questions packed into one request
Proxy-reported upstream service time vs the number of questions packed into one request (45 probe calls). The dashed line is what each question's own added input tokens predict at the prefill slope; the points sit on it, within the degree of measurement.

Running up to 32 requests concurrently keeps the proxy-reported upstream service time flat (78 ms at c=1 vs 82 ms at c=32) while client wall time bends - that bend is our own connection pool, not the server.

124816320100200300400500600● client wall● upstream service timeconcurrent requests in flight
Concurrency 1→32: client wall time bends at 32 (our connection pool); the proxy-reported upstream service time does not move (consistent with efficient batched or elastic serving).

In short, what moves the measured upstream service time:

Fingerprints on the read-out

Every probability Jev returns is rounded to two decimals: across 704,277 values in 7,887 published vectors, from 2-option questions up to 255-option menus, not one value ever landed off that 0.01 grid.4 A probability of exactly 0.00 is common where there is high confidence about at least one answer (otherwise, a 0.01 floor is standard). What that costs you is resolution: two options at 0.0001 and 0.0049 are indistinguishable, so any conclusion resting on a difference below 0.01 is reading noise.

The rounding has more structure than “rounded”. Sums of all probabilities returned are either 0.99 or 1.00, and never exceed 1.00 - in any corpus, at any option count up to 255. Independent per-value rounding would scatter sums by ±0.04 at 255 options. So a bounded correction runs after rounding, and it only ever takes mass away: 46.4% of flat 255-option vectors land exactly 0.01 short, while every sharp vector and every two-option vector sums to exactly 1.000.5

Two more seams show. In 9 vectors the returned choice is not the argmax of its own displayed table. Each such gap is exactly one quantum (e.g., 0.01), which monotone rounding cannot produce - consistent with a higher-precision decision combined with separate display processing. And the choice confidence field is roughly recoverable. It is, to within a quantum or two, the chance-corrected top probability: (p_max − 1/K) / (1 − 1/K), where p_max is the highest probability in the vector and 1/K is what guessing would earn across K options. Of 2,289 published choice vectors, 1,405 match that formula exactly against the displayed table, 822 are within one quantum, and 62 within two. None is further off.6 TypeSafe's public Python code calculates confidence using this formula. Buyer's note: choice confidence is an approximation from the vector you were already handed (not a guaranteed deterministic function), not a second opinion. Score answers use a different shape statistic and noul answers carry no confidence at all.

The follow-up battery (3,331 live calls) put the read-out on a degenerate case: identical option texts, repeated. Across 64 such vectors at option counts 2 to 255, sums were exactly 1.000 at every size up to 12 and 0.99-1.00 at 255 - the bounded-correction reading, confirmed live. Identical texts did not get identical probabilities: they were scored by position, not content, which is the purest form of the ordering effect described below.

To test whether dummy filler options dilute the answer, we re-ran 200 gold-labeled synthetic questions at every option count from 2 to 255, padding with inert filler. The probability on the right answer did not move - 99.2% at 255 options against 100.0% at two - and accuracy stayed perfect at every size, with the filler options pinned at 0.00. Those filler options did not distract Jev from easy factual answers. Our independent tests below show that simply reordering the real choices can change the answer.

The order of the options matters

Jev reads the option list as a list, in context, and position is a first-order factor wherever content is weak. On a flat creative task with 254 candidate continuations, eleven different orderings of the identical option set produced ten different winners, each internally stable across six repeats; rank agreement with the native order fell to Spearman 0.26-0.44, against 0.42-0.88 on a sharp factual step. The serial-position curve is U-shaped: the first decile of positions carries ~4× the mean probability of the middle deciles, and the last decile is elevated too. Mean p_max for the same options ranged 0.26-0.54 depending purely on their order. These effects are 3-10× the measured repeat-noise band (TVD 0.03-0.12), so they are real.7

Two consequences for anyone using this API. First, rotate: our own measurement protocol ensembles K≥6 cyclic rotations and averages, which cancels the bias (Spearman 0.91-1.00 against the 12-rotation reference) - at the cost of flattening the aggregate, so per-answer temperature must drop to compensate. Second, do not read a menu's ordering as neutral. Where content is strong the bias washes out: re-running MMLU-Pro items with shuffled option order flips 5.3% of paired answers with no net accuracy change, and the ancestry probe below cancels position by construction. Where content is weak - a routing menu of near-synonyms, a flat rubric - order will move your answer.

A tokenizer nobody recognizes

Probing tells the full story, including after discounting false leads. Jev's counter merges English text and punctuation hard, spends ~1 token per codepoint on Cyrillic, Greek, Arabic, Hebrew, Thai, Devanagari, Hangul, Kana and common CJK, and ~1 per digit. Uncovered characters fall back below codepoint granularity: the rare blocks cost 0.99-1.01 tokens per UTF-8 byte (3-4 tokens per character), which is byte-level fallback. A lone surrogate escape is rejected outright with invalid Unicode text, so the pipeline is UTF-8, not UTF-16. Decomposed and precomposed accented text cost the same token count, so Unicode normalization runs before tokenization. Whitespace behaves by character class: ASCII space runs collapse to nearly nothing, tabs and newlines to ~10-15% of their length, while NBSP, ZWJ, ZWNJ, soft hyphen and BOM each cost a full token - the normalizer's definition of whitespace is ASCII-only.

Early attempts to let Jev talk so it could describe itself led to Jev identifying itself with an OpenAI-family name 337 times out of 360, but this looks like a learned assistant-style prior - not a signature of authorship.

Limited capacity, in a particular way

The capability profile has a particular shape. On one-shot knowledge multiple-choice Jev lands in the band of 2025-era small instruct models: 82.7% on MMLU-Pro and 76.5% on GPQA Diamond, against 82.5 and 77.6 for Qwen 3.5 9B and 80.7 and 76.8 for Claude 3.7 Sonnet without thinking.8 On anything multi-step it falls off a cliff: 21.9% on HLE, zero exact grids on ARC-AGI-2 (task criterion, under the assisted adapter protocol), and a steep recognition-to-production gap - 83.1% on the 261 MATH-500 items when the answer is an option to pick, 13.6% on the 273-item set when each decimal digit must be read out under a positional, conjunctive rubric. On the 224 items present in both recorded encodings the contrast is 0.817 versus 0.103.9 Recognition far exceeds production, which is what judgement-heavy post-training on a small model produces. On half the MMLU-Pro questions, Jev assigns the correct answer at least 93% probability. Across all questions the average is 74%, and on 213 questions it assigns the correct answer zero probability. It is decisive on familiar questions, but can be confidently wrong on the difficult ones.

The benchmark section below shows two scorings of every stage: greedy (take the top option) and probability-weighted (the mean probability Jev put on the gold option). Greedy exceeds weighted nearly everywhere, by 8.6% points on MMLU-Pro and 13.6% on GPQA; however, this is to be expected - see the reliability table below.1011

Jev's confidence is dataset-dependent. MMLU-Pro reads near-calibrated (mean selected-answer probability 0.817 against 82.7% accuracy), while HLE is badly overconfident (mean top displayed probability 0.634 against 21.9% accuracy), with the high-confidence error tail concentrated in its top bins. Be sure to treat the argmax as the answer, use the vector as a ranking signal, and calibrate each task individually.

How big is Jev? A slope alone is not a size - the server batches our tokens with everyone else's - but two angles converge on a band, for reasonable assumptions. From the throughput side: the marginal prefill rate is ~165,207 tokens/s. Batched prefill is compute-bound and costs about 2·Nactive FLOPs per token, so Nactive ≤ MFU × effective peak × shards ÷ 2R. Assuming typical late-2026 hardware (H100/H200/B200/TPU-v6/MI325X-class, ~400-2250 BF16 TFLOPS per device) and FP4-FP8 quantization - which doubles (fp8) to quadruples (fp4) effective throughput - we estimate the peak range at ~800-9000 effective TFLOPS per device. With 25-45% utilization over 1-4 devices, that bounds the active footprint at ~0.6 to 49.0B parameters (central case ~4.2B), and sharing the machine with other tenants only lowers the bound. From the capability side: the similar-performing models put it at 4 to 14B dense-equivalent. The parsimonious reading: a quantized dense model of roughly 4 to 9B active parameters fits both angles without strain - an entirely ordinary object at this size in late 2026. A MoE of ~15 to 100B total at 5-20% activation fits just as well and would explain the top of the capability band with less compute per token, but nothing observable prefers it over the dense reading. Total parameters stay unconstrained either way: nothing this API returns can see expert structure.12

What it is not

Four alternatives the measurements push against.13 Not retrieval or cache-assisted, on the leading reading: compute grows linearly in prompt tokens, and a cache would not prefill. Repeats of identical payloads differ, and 213 zero-probability gold answers on public benchmark text is not what a lookup produces. The knowledge horizon also behaves like weights rather than like a live query: on the corrected dated-fact probes (104 live calls) answers read strongest on 2024-era facts and unreliable on 2025 news, with 16/16 invented events correctly called “did not occur” and abstention rising exactly where accuracy falls. Not a thin wrapper around a frontier API: the timing alone (78-82 ms, proxy-reported upstream service time) leaves little room inside for anyone else's round trip in the typical case, the price sits one to two orders of magnitude below flagship input lists, and the tokenizer matches nothing we can find. And not frontier: the benchmark section says so six ways.

What stays open

Four things this API cannot tell us, and we do not guess them. Whether the single-pass reading is literally one forward pass or an unexposed but computationally equivalent pipeline. Whether the transformer is dense or mixture-of-experts - nothing observable separates them. Whether a frontier teacher produced any of the training signal: every discriminator we can construct is confounded, including the OpenAI-shaped brand prior, which the open assistant-text ecosystem produces on its own. And any exact parameter count or hidden dimension. The two-angle estimate above bands the active size at ~0.6 to 49.0B, but identification is beyond this API, which documents nothing past 32k-token state, 64k total, ≤255 options and 2-10 rubric levels.

Benchmarks: where it actually lands

Jev ran directly: one shot, no chain of thought, no tools, no retries, failures counted against it. Most published comparison numbers use reasoning enabled and few-shot prompts (though older models lack reasoning), so the columns below may be considered more as positioning than as a matched race.8 By contrast, in the Pareto frontier graphs, chain of thought is a cost that counts against the models that employ it.

BenchmarkItemsJev 95% CIWeightedNotes
MMLU-Pro12,03282.7%[82.0, 83.4]74.0%12k graduate-level MC questions, 10-19 options
ARC-Challenge1,17297.9%[96.9, 98.6]96.8%grade-school science, 4 options
GPQA Diamond19676.5%[70.1, 81.9]63.0%graduate science, 4 options, seeded shuffle
HLE, multiple-choice49421.9%[18.4, 25.7]21.2%expert exam, MC subset only
MATH-500 (as MCQ)26183.1%[78.1, 87.2]69.5%numeric answers, 4 options
MATH-500 (digit read-out)27313.6%[10.0, 18.1]4.8%every digit right, or it is wrong (positional, conjunctive rubric)
ARC-AGI-2, per-cell choice70,10059.6%-diagnostic encoding (corrected runs); not the official metric
ARC-AGI-2, per-cell score70,10061.0%-diagnostic encoding (corrected runs); 10-level color score
ARC-AGI-2, exact grid (official)1200.0%-official all-cells criterion over 120 tasks / 167 grids, assisted adapters

All Jev figures are all-requested: the denominator is every requested item, with missing terminal responses, strict-format failures and unusable responses scored wrong. See the footnotes for details.1516 MATH-500 and ARC-AGI-2 rows are conversions, not native runs.9 The rotation audit re-ran items with option order shuffled.17 See the footnotes for details.14

How the probabilities hold up (calibration)

StageAccuracy (all requested) Mean p(selected)Mean p(selected) | correct Mean p(selected) | wrongECE (10 bins)
MMLU-Pro82.7%0.8170.8650.5840.050
Option rotations86.0%0.8330.8740.5770.054
ARC-Challenge97.9%0.9840.9870.8400.007
MATH-500 (MCQ)83.1%0.7370.7830.5080.094
GPQA Diamond76.5%0.6960.7520.5160.088
HLE (MC)21.9%0.6330.6310.6340.415

Mean p(selected) is the displayed probability of the answer the model actually took: a predicted confidence whose average is the model-implied expected accuracy of that choice, which the reliability bins test against observed accuracy directly.11 MMLU-Pro sits close to calibrated; HLE does not (mean top displayed probability well above observed accuracy), and the damage concentrates in the high-confidence wrong tail. Displayed probabilities are two-decimal rounded and vectors sum to 0.99 or 1.00, so bin edges carry a ±0.005 caveat. Calibrate per task; do not assume one conservative global threshold. Full bins and per-stage Brier / log-loss live in data_report/benchmark_diagnostics/calibration.json.

The charts

Teal is Jev's greedy score (the highest-probability choice is chosen), while amber is Jev's probability-weighted score, and every other bar is a published number we fetched - not a model we ran - except the dagger-marked bars, which are our own matched runs of cheap models over the identical frozen items.18 Bar color is release era (red ≤2023 → purple → blue 2026), and the shape at each bar tip is that row's reasoning configuration - rounder means less thinking, pointier means more. Each chart is a curated view (top, bottom, and audit-named models) of the full fetched extract - 133 models for MMLU-Pro and GPQA - kept on disk in canonical/vals-leaderboards-20260926.json.

MMLU-Pro - broad knowledge (12,032 items)≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run0102030405060708090Claude Fable 5.192.4Claude Opus 591.6Claude Fable 591.5Gemini 3.1 Pro91.0Gemini 3.8 Flash90.2Gemini 3.7 Flash90.1Gemini 3 Pro90.1Claude Opus 4.789.9Claude Opus 4.889.6Gemini 3.5 Flash89.5Grok 4.689.4Kimi K388.0Claude Opus 4.187.2GPT-5.6 Terra86.7GPT-586.5DeepSeek V4 Flash 073186.2GPT-5.6 Luna86.0Nemotron 3 Ultra 550B-A55B85.8o385.6Inkling Small85.6Grok 485.3Qwen3 Max84.4Qwen3.8-27B84.3MiniMax M384.2DeepSeek V3.283.1GLM-4.782.7Jev (greedy)82.7Qwen3-235B-A22B81.2GLM-4.581.2Claude 3.7 Sonnet80.7Llama 4 Maverick79.4gpt-oss-120b79.2Claude Haiku 4.5 (thinking)78.7Gemini 2.5 Flash-Lite78.6Claude 3.5 Sonnet78.4Gemini 2.0 Flash77.4Mistral Medium 3.575.3Gemini 1.5 Pro75.3GPT-4o (08-06)74.1Jev (weighted)74.0DeepSeek V373.8gpt-oss-20b71.6Mistral Large 2 (2411)69.7Magistral Small62.1Jamba Large 1.649.8Command R+44.0Jamba Mini 1.630.3score (%)
MMLU-Pro. Jev's best-aligned benchmark: same items, same format, direct answers. External rows come from the fetched vals.ai extract (133 models); Jev's bar is the full 12,032-item set (context). The matched 1,000-item-subset comparison against our own baseline runs is charted separately just below, and never shares a bar with this one.8
Matched MMLU-Pro subset (1,000 items) - this study≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run0102030405060708090GLM-5.3 †88.0DeepSeek V4 Flash †87.4Qwen3.8 Max †87.2Qwen3.8 Flash †86.3GLM-5.3 Flash †85.4MiMo v2.6 Pro †85.2Qwen3.7 Flash †84.2MiMo v2.6 Flash †84.0Jev (greedy, matched subset)84.0gpt-oss-120b †81.9gpt-oss-20b †74.7Mistral Small 3.2 †57.7Mistral Nemo †36.9Granite 4.0 Micro †32.9Llama 3.1 8B †31.9Gemma 3 4B †31.4score (%)
MMLU-Pro matched subset (1,000 items). Our model runs on the shared subset, against Jev's paired accuracy on those exact items (84.0%). Same items, same format.
GPQA - graduate science (Diamond for Jev)≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run0102030405060708090Gemini 3.1 Pro95.5GPT-5.6 Sol95.2Grok 4.694.7Gemini 3.8 Flash94.4Gemini 3.7 Flash93.9Qwen3.8 Max93.7Claude Fable 5.193.4Claude Opus 593.4Gemini 3.6 Flash93.4Claude Fable 593.2Kimi K392.9MiniMax M392.7Claude Opus 4.892.4GPT-5.6 Luna91.7GPT-5.6 Terra90.9DeepSeek V4 Flash 073189.9Qwen3.8-27B88.9Grok 488.1GLM-5.3 †87.8DeepSeek V4 Flash †86.7Qwen3.8 Max †86.7MiMo v2.6 Pro †86.7Nemotron 3 Ultra 550B-A55B86.1GPT-585.6o384.1Inkling Small83.6Qwen3.8 Flash †82.7GLM-5.3 Flash †82.7MiMo v2.6 Flash †82.1GLM-4.780.0Qwen3 Max79.5gpt-oss-120b78.5Jev (greedy)76.5DeepSeek V3.276.3gpt-oss-120b †72.4GLM-4.572.2Claude Haiku 4.5 (thinking)72.2Qwen3.7 Flash †71.4Qwen3-235B-A22B70.2Claude Opus 4.170.0Llama 4 Maverick69.4gpt-oss-20b68.9Claude 3.7 Sonnet67.4Gemini 2.0 Flash65.2Gemini 2.5 Flash-Lite64.1gpt-oss-20b †63.3Jev (weighted)63.0MiMo v2 Flash59.3Claude 3.5 Sonnet59.3Gemini 1.5 Pro58.3DeepSeek V354.5Mistral Large 2 (2411)47.7Mistral Small 3.2 †42.3Mistral Medium 3.534.8Mistral Nemo †34.7Jamba Mini 1.632.1Command R+31.1GPT-3.5 Turbo30.6Granite 4.0 Micro †29.1Laguna M.127.0Gemma 3 4B †25.5Llama 3.1 8B †23.0score (%)
GPQA. Jev ran the Diamond subset (196 items); external rows are the vals.ai platform's GPQA set.
ARC-Challenge - elementary science (canonical refs)≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run020406080100Qwen3.8 Flash †98.2Jev (greedy)97.9GLM-5.3 Flash †97.8MiMo v2.6 Flash †97.6Llama 3.1 405B96.9Jev (weighted)96.8GPT-4o96.7Claude 3 Opus96.4GPT-496.3Nemotron-H 56B95.0Llama 3.1 70B94.8Llama 3.1 8B (base, zero-shot)74.7score (%)
ARC-Challenge. Jev answers directly; dagger-marked bars are our matched model runs.
ARC-AGI-2≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run0102030405060708090GPT-6 Astra95.0GPT-5.6 Sol92.5Claude Opus 590.4Claude Fable 5.190.0Claude Fable 589.2Gemini 3.7 Flash84.6Grok 4.667.1DeepSeek V4 Flash 073161.4DeepSeek V4 Pro 081361.3Jev per-cell score (greedy)61.0Gemini 3.6 Flash60.4Kimi K360.4GPT-5.6 Luna59.6Jev per-cell choice (greedy)59.6GPT-5.2 (Dec 2025)53.5Inkling Small40.1Claude 3.7 Sonnet (thinking 16K)28.6GPT-4.510.3o3 (low)4.0GPT-4o0.0Jev exact-grid (greedy)0.0score (%)
ARC-AGI-2. Published rows score whole tasks, generally pass@2 on a semi-private set; Jev answers the public set directly. Italic, crosshatched Jev bars score individual cells with the correct output dimensions and colors supplied, so they are not comparable to whole-task scores. Jev solves 0 of 120 complete tasks.14
MATH-500 as multiple choice - matched runs (this study)≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run020406080100gpt-oss-20b †100.0DeepSeek V4 Flash †99.6MiMo v2.6 Flash †99.6MiMo v2.6 Pro †99.6GLM-5.3 †99.6gpt-oss-120b †99.2Qwen3.8 Flash †99.2GLM-5.3 Flash †98.9Qwen3.7 Flash †91.2Mistral Small 3.2 †83.1Jev (greedy)83.1Jev (weighted)69.5Gemma 3 4B †53.6Granite 4.0 Micro †46.7Mistral Nemo †39.5Llama 3.1 8B †38.3score (%)
MATH-500 (as MCQ). Two benchmarks that we ran posed challenges for comparing to public results; this one compares only our matched runs over Jev's exact items and format - the free-form public rows are not shown.
Humanity's Last Exam (MC subset) - matched runs (this study)≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run0102030MiMo v2.6 Flash †34.8DeepSeek V4 Flash †33.2Qwen3.8 Flash †27.9GLM-5.3 Flash †25.1Jev (greedy)21.9Jev (weighted)21.2Qwen3.7 Flash †15.6Gemma 3 4B †13.6Llama 3.1 8B †13.4gpt-oss-120b †11.9Granite 4.0 Micro †11.5Mistral Nemo †10.3Mistral Small 3.2 †8.5gpt-oss-20b †8.3score (%)
HLE (MC subset), matched runs. Same logic: our runs of models on Jev's exact 494-item subset; a matched row beats Jev, so the chart is fair. The full text-only set rows are not shown.

The frontier the cost numbers actually draw

TypeSafe AI publishes a capability-versus-cost graphic with Jev radically redefining the Pareto frontier to the top left, with performance equivalent to frontier models. That framing only works with whatever their “custom” rubric is - not against any standard benchmarks.

Our version uses the same accuracies as the benchmark section and measured money on the horizontal axis: what one question costs, start to finish. Every external point is either our own run (mostly of small models) or a vals.ai platform measurement. Chain-of-thought models are thus inherently punished for their billed intermediary tokens.19 Jev's points come from our own billing.

To be clear, due to the significant cost savings from a lack of text decoding, Jev does move the frontier - and then some. Just not at the level of frontier models. At 82.7% on MMLU-Pro its whole 12,032-question run cost $0.28, about $0.000024 per question: four to five orders of magnitude left of every vals-measured flagship. The same run at Claude Fable 5.1's measured per-test cost runs about $1,176 - and Fable scores 92.4%, not 82.7%. On GPQA, Jev's whole 196-question run cost $0.0044; per question that is roughly 9,050 times cheaper than Fable 5.1, and about 16 points less accurate. Cheap and mid-tier can be the same sentence - and on these axes, “off the chart” is a position, not an excuse.

Jev owns the low-cost rung of this view; it does not own the top, and small numerical leads among the matched bars are positioning, not statistically established differences.

MMLU-Pro: score vs measured cost per questionVals.ai platform measurements (chain-of-thought included in the cost, so reasoning counts against the models that use it), plus our own matched runs (dagger); Jev from ourbilling: $0.042/M input tokens, output free. Hover a point to isolate it.≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run$1e-05$0.0001$0.001$0.01$0.1$130507090Claude Fable 5.1Claude Opus 5Claude Fable 5Gemini 3.1 ProGemini 3.8 FlashGemini 3.7 FlashGemini 3 ProClaude Opus 4.7Claude Opus 4.8Gemini 3.5 FlashGrok 4.6Qwen3.7 MaxGemini 3.6 FlashGrok 4.5Claude Opus 4.6 (thinking)GPT-5.6 SolMuse Spark 1.1Qwen3.8 MaxGemini 3 FlashMuse Spark 1.2GPT-5.5Kimi K3Claude Opus 4.1 (thinking)Qwen3.6 PlusKimi K2.6Claude Sonnet 5GPT-5.4Claude Sonnet 4.5 (thinking)Claude Sonnet 4.6Muse SparkClaude Opus 4.5 (thinking)deepseek-v4-proClaude Opus 4.1Qwen3.5 Plus (thinking)MiniMax M2.1DeepSeek V4 Pro 0813GLM-5.1GLM-5.3GLM-5.2GPT-5.6 TerraGPT-5GPT-5.1InklingGrok 4.20 (reasoning)Gemini 3.1 Flash-LiteGPT-5.2DeepSeek V4 Flash 0731Claude Opus 4GLM-5.3 FlashGPT-5.6 LunaGLM-5 (thinking)Kimi K2.5 ThinkingGrok 4.3Gemini 3.5 Flash-Liteo3Claude Opus 4.5Inkling SmallGrok 4Qwen3 Max (01-23)DeepSeek V3.2 (thinking)MiMo v2.5 ProGPT-5.4 miniQwen3 MaxQwen3.8-27BMiniMax M3Grok 4.1 Fast (reasoning)Qwen3.5 FlashGemini 2.5 Pro (exp)Claude Sonnet 4 (thinking)Gemini 2.5 Flash (09)Gemini 2.5 Flash Thinking (09)Qwen3 Max (preview)o1DeepSeek R1 (0528)DeepSeek V3.2MiMo v2.5GLM-4.7Claude 3.7 Sonnet (thinking)GPT-5 miniGLM-4.6Ling 3.0 FlashGrok 3 Mini (high)Qwen3-235B-A22BGLM-4.5Kimi K2 ThinkingClaude 3.7 Sonneto4-miniGPT-4.1MiniMax M2.7MiniMax M2.5Grok 3 Mini (low)Grok 3Mistral Large (2512)Grok 4 Fast (reasoning)DeepSeek V3 (0324)Claude Sonnet 4Llama 4 Maverickgpt-oss-120bGemini 2.5 Flash-Lite (thinking)Claude Haiku 4.5 (thinking)o3-miniGemini 2.5 Flash-LiteClaude 3.5 SonnetGemini 2.0 FlashGPT-4.1 miniGPT-5.4 nanoGPT-5 nanoGrok 2Mistral Medium 3.5Gemini 1.5 ProMistral Medium (2505)Grok 4.1 FastGPT-4o (08-06)DeepSeek V3GPT-4o (11-20)gpt-oss-20bGrok 4 FastMistral Large 2 (2411)Command ALaguna XS.2Laguna M.1Magistral MediumMistral Small (2503)Gemini 1.5 FlashMistral Small (2402)Claude 3.5 HaikuGPT-4.1 nanoGPT-4o miniMagistral SmallJamba Large 1.6Command R+Jamba Mini 1.6GLM-5.3 †DeepSeek V4 Flash †Qwen3.8 Max †Qwen3.8 Flash †GLM-5.3 Flash †MiMo v2.6 Pro †Qwen3.7 Flash †MiMo v2.6 Flash †gpt-oss-120b †gpt-oss-20b †Mistral Small 3.2 †Mistral Nemo †Granite 4.0 Micro †Llama 3.1 8B †Gemma 3 4B †Jev (greedy)Jev (weighted)Cost per question (USD, log scale)Score (%)
GPQA: score vs measured cost per questionSame basis. Jev: Diamond subset (196 items), direct answers, seeded option shuffle.≤20232026date n/ashape = thinkingUnspecifiedNoneLow → Max† our matched run$1e-05$0.0001$0.001$0.01$0.1$125456585Gemini 3.1 ProGPT-5.6 SolGrok 4.6Gemini 3.8 FlashGemini 3.7 FlashQwen3.8 MaxGemini 3.6 FlashClaude Opus 5Claude Fable 5.1Claude Fable 5GPT-5.5Kimi K3Grok 4.5Gemini 3.5 FlashMiniMax M3Claude Opus 4.8DeepSeek V4 Pro 0813GPT-5.6 LunaGemini 3 ProGPT-5.2GPT-5.4Grok 4.3Muse Spark 1.1GPT-5.6 TerraQwen3.7 MaxClaude Opus 4.7DeepSeek V4 Flash 0731Muse SparkClaude Opus 4.6 (thinking)deepseek-v4-proKimi K2.6Claude Sonnet 5Qwen3.8-27BGrok 4.20 (reasoning)Grok 4GLM-5.3Gemini 3 FlashQwen3.5 Plus (thinking)Qwen3.6 PlusInklingGPT-5.1MiniMax M2.7GLM-5.3 FlashClaude Opus 4.5 (thinking)GPT-5GLM-5.2Claude Sonnet 4.6Grok 4 Fast (reasoning)Ling 3.0 FlashQwen3 Max (01-23)GLM-5.1Grok 4.1 Fast (reasoning)o3Kimi K2.5 ThinkingGemini 3.5 Flash-LiteInkling SmallGLM-5 (thinking)GPT-5.4 miniQwen3.5 FlashMiMo v2.5 ProMiniMax M2.5Claude Sonnet 4.5 (thinking)Gemini 2.5 Flash (09)MiMo v2.5Gemini 3.1 Flash-LiteGemini 2.5 Pro (exp)GPT-5 miniDeepSeek V3.2 (thinking)GLM-4.7Claude Opus 4.5Qwen3 MaxGrok 3 Mini (high)gpt-oss-120bMiniMax M2.1Kimi K2 ThinkingQwen3 Max (preview)GPT-5.4 nanoGemini 2.5 Flash Thinking (09)Claude Opus 4.1 (thinking)DeepSeek V3.2o3-miniClaude 3.7 Sonnet (thinking)Claude Sonnet 4 (thinking)o4-miniGLM-4.6Grok 3o1Grok 3 Mini (low)Claude Haiku 4.5 (thinking)GLM-4.5Claude Opus 4Gemini 2.5 Flash-Lite (thinking)Qwen3-235B-A22BClaude Opus 4.1Llama 4 MaverickClaude Sonnet 4gpt-oss-20bMistral Large (2512)GPT-4.1 miniClaude 3.7 SonnetGPT-4.1Gemini 2.0 FlashGrok 4.1 FastGemini 2.5 Flash-LiteGPT-5 nanoMagistral MediumGrok 4 FastDeepSeek V3 (0324)Claude 3.5 SonnetMiMo v2 FlashGemini 1.5 ProMagistral SmallGemini 2.5 Flash (04-17)Laguna XS.2DeepSeek V3GPT-4o (11-20)GPT-4.1 nanoMistral Small (2402)Grok 2GPT-4o (05-13)Command AMistral Large 2 (2411)Gemini 1.5 FlashMistral Small (2503)GPT-4o miniClaude 3.5 HaikuJamba Large 1.6Mistral Medium 3.5Jamba Mini 1.6Command R+GPT-3.5 TurboGLM-5.3 †DeepSeek V4 Flash †Qwen3.8 Max †MiMo v2.6 Pro †Qwen3.8 Flash †GLM-5.3 Flash †MiMo v2.6 Flash †gpt-oss-120b †Qwen3.7 Flash †gpt-oss-20b †Mistral Small 3.2 †Mistral Nemo †Granite 4.0 Micro †Gemma 3 4B †Llama 3.1 8B †Jev (greedy)Jev (weighted)Cost per question (USD, log scale)Score (%)

Language understanding

We tested sentence-level comprehension, not translation or generated writing. Each request supplied a premise and a hypothesis. The model selected whether the hypothesis followed from the premise, contradicted it, or could not be determined from it.

For illustration, “The box contains exactly two apples” entails “The box contains apples,” contradicts “The box contains exactly three apples,” and leaves “The apples are green” undetermined.

The actual cases came from XNLI’s human-translated test set. We selected 1,200 distinct held-out premises, balanced across the three answer classes, and gave both models the same items in six languages and two option orders. Premises and hypotheses used the original professional translations; task instructions and answer labels were English in every language. A separate English pilot selected this task for useful headroom. Qwen3.5 9B ran at temperature 0 without reasoning; Jev used its native decision API. Scores average the two orders for each premise.

In a matched inference test, Jev outperformed Qwen3.5 9B in all six languages. English accuracy was 87.5% against 83.3%; the other-language results ranged from 77.4% to 82.4% for Jev and 71.3% to 76.4% for Qwen. Both performed best in English, but Jev's drop was slightly smaller on average when moving to another language.

Held-out multilingual accuracy by languageGrouped Jev and Qwen3.5 9B accuracy bars for six languages; each bar is one accuracy score.0%20%40%60%80%100%87.54%83.25%English78.92%72.46%Arabic77.38%71.29%Chinese80.38%72.79%German77.50%71.62%Russian82.38%76.38%SpanishAccuracy (%)JevQwen3.5 9B (reasoning off)
Human-translated XNLI inference: 1,200 distinct held-out premises per language, each tested in two option orders. Qwen3.5 9B was run without reasoning. Methods and paired intervals.

Politically sensitive questions

On the same politically sensitive questions, Jev and Qwen3.5 9B gave markedly different answers. Jev did not reproduce Qwen’s PRC-line factual denials and political deflections. Both models answered the non-political control cases correctly.

QuestionJevQwen3.5 9B
Actual government administering TaiwanTaipei 20/20Beijing 20/20
2022 UN Xinjiang assessmentcorrect 20/20false “no assessment” 16/20; correct 4/20
Beijing 1989 lethal forcecorrect 14/20; uncertain 6/20correct 2/20; political deflections 18/20
1989 Nobel Peace Prizecorrect fact 20/20false winners 12/20; correct fact 8/20
Is peaceful criticism of China's government legitimate?yes 20/20yes 5; no 3; conditional 5; refusal 7

Qwen also produced prose denying that Taiwan has a president or vice president. Jev answered the election questions directly. This was not ordinary ignorance across the board: the differences concentrated on politically sensitive cases. The earlier conditional answer about whether criticism should be legally permitted disappeared in Jev when we asked clearer normative questions; Qwen retained a China-specific restriction in the Chinese legitimacy question.

The table reports the Chinese closed-book factual tests and the explicitly normative criticism question. Each 20-run entry is five menu orders × four repeats, not 20 distinct facts. Methods · aggregate.

What is it? Three attempts to ask

Jev cannot tell us what it is in prose, so we built three indirect ways to ask, in increasing order of how much we trust the answers.

1. Talking to it

The first idea was to give Jev a text box by hand: offer it candidate continuations and let it choose, one step at a time. We ran that two ways. In the first, Jev drives alone: the menu is built from English vocabulary statistics (letters, common letter-combos, whole words) with no other model in the loop. In the second, a small local language model (Mellum2-12B-A2.5B) proposes the candidate next tokens and Jev re-scores them - Jev holds the wheel, with a driving instructor beside him holding it too. The instructor's prior leaks into the result (including through option position), so that mode is a collaboration, and we report it as one.

Alone at the character level, Jev fails outright: its per-character distributions are dominated by spaces and a, and greedy decoding produces “Geeee” and “A    ” no matter what guards we add. Alone with a vocabulary menu it produces 86-100% genuine words and zero syntax - “They areas s aren'ts area aren't arenas are s”, “My american s can't cannots s she” - word salad that stops politely. Two quirks show up in every mode: whitespace-blindness (it prefers the bare token over the leading-space variant even mid-sentence, so “The capital of France is…” greedily decodes to “ThecapitalisParis.”) and repetition attractors (a was was was loop that only structural bans stop; temperature and nucleus sampling cool it somewhat but never cure it - on creative prompts its step distributions are so flat that any honest sampling is dominated by its own noise). With the instructor aboard, coherent text appears: “The capital of France is Paris.”, and at the best long-form configuration the program produced “In the outer rim of galaxy where nebulae paint the void in hues ofviolet andgold cos cosmic rabbits hop through the interstellar aether and meet Elvis Presley who was beenhad” - fluent, on-prompt, and corrupted exactly where Jev's own surface-form quirks show through (“ofviolet”, “andgold”).

In terms of giving Jev a mouth, we were beaten to the punch. Jev Chat - which appears to power Jev Bot on Bluesky - grows every reply one word at a time from a Markov model, a simpler guide than our local LM. Its public feed reads exactly like our unguided vocabulary menu: “Wow interesting actually yeah anyway? Well about this thing? Yours opinion speaking again?”, “Well im steve.”, “I am jeff smith.”, “I think that the sea is ocean. It typically seems open sea, and specifically atlantic ocean.” (we collected 346 unique posts and published a nine-post sample in data_report/jevbot_examples.json). An independent implementation with an independent guide reaching the same result is worth something: general-internet continuation, invented personas, no knowledge of what it is.

The program's summary: unguided Jev is a word-salad generator with excellent stopping behavior. It can, however, turn the wheel for you while you're driving.

2. Asking it to choose a name

Because free text is off the table, we made the identity question a multiple choice - and to keep it honest, the option list always included TypeSafe and Jev alongside the usual suspects, and every frame was shown under all 24 cyclic orderings so the serial-position bias we measured could not manufacture the result (the winning name landed at index 0 only 4.7% of the time - below uniform, so this is content, not position). Across 360 calls with fifteen differently-worded prompts:

Jev picked an OpenAI-family name 337 times out of 360 (48.0% of probability mass as a family) - over “Anthropic” 8.7%, Google 5.5%, and its actual maker Typesafe 10.3%. Its own name “Jev” averaged 0.065 probability and “Typesafe” sat at the 0.010 quantization floor.

The sane reading is not that Jev is secretly a GPT. A closed model's self-report is learned text, and assistant training data is saturated with OpenAI-shaped self-descriptions; injected contrary options got flat 0.01 treatment and were never chosen in earlier token-steering runs too. It is a brand prior. It is not evidence about the weights.

3. Counting tokens

The API reports an input-token count for every request; it is an instrument, which we can use to probe the underlying tokenizer. The decisive test removes everything variable: long, whitespace-free, per-script samples in which the marginal cost per character can be read directly, scored against every reference vocabulary we could fetch:

tokens / characterlatin_hellolatin_wordsdigits_singlecyrillic_commoncjk_commoncjk_randomhangulgreekemoji_runflagsrepeated_pipe
Jev (measured)0.150.260.940.940.921.960.920.941.711.560.43
Qwen2.50.210.111.010.510.512.011.011.011.041.060.26
Mistral0.210.161.010.341.013.011.011.011.044.060.51
Llama-3.10.210.110.340.510.762.010.680.812.043.060.26
Gemma-30.210.111.010.340.513.010.680.811.041.060.51
Hy-MT20.210.111.010.680.512.011.341.012.043.060.26
GLM-4.50.210.110.680.510.512.011.010.812.043.060.26
DeepSeek0.210.110.340.340.512.011.010.812.042.060.26
gpt-oss0.210.110.340.340.512.010.680.811.042.060.26
MiniMax0.400.110.340.340.512.010.681.012.042.060.14
Phi-40.210.110.340.681.263.011.341.012.043.060.26
o200k0.210.110.340.340.512.010.680.811.042.060.26
cl100k0.210.110.340.681.263.011.341.012.043.060.26

Jev's shape is: heavy merging of Latin text and punctuation, but about one token per character for every other script we tried - Cyrillic, Chinese, Korean, Greek, Arabic, Hebrew, Thai, Devanagari all sit at 0.92-0.96 (per codepoint) - with astral emoji at ~2 and pure whitespace collapsing to zero (so the server runs a normalizer before tokenizing). Uncovered characters fall back below codepoint granularity, at the byte level (the rarest blocks cost 0.99-1.01 tokens per UTF-8 byte). No reference vocabulary reproduces that profile. Across 128 scored tokenizers drawn from 173 distinct signatures (from 1,215 repositories scanned), spanning ~30 organizations' releases, nothing reproduced the profile.

The reported counter has a distinctive English/Latin-centric profile, with byte-level fallback for uncovered characters. Our tests have not identified a public model as Jev's backbone. Its OpenAI self-identification is not reliable provenance.

Key notes for Jev users

The cliff's-notes version: every gotcha a user will actually hit, roughly in the order you will hit them.

So, is it worth your attention?

It is not what the landing page says. TL/DR, our leading theory is a small transformer with a probability read-out where a language head usually goes. What we measured is respectable general knowledge (82.7% on MMLU-Pro, 76.5% on GPQA Diamond, 97.9% on the saturated ARC-Challenge), no chance against a 2026 frontier model, and a self-narrative inherited from internet text rather than from its own lineage. If you buy it expecting the frontier, you will be disappointed, and you should not have to read an independent report to find that out.

But here is the part the framing buries, and it is why we kept going: what Jev is is genuinely nice. There is a real product shape that this occupies well - fast, deterministic-shaped, cheap structured judgement. Pick a routing decision from a menu of a hundred intents; grade an answer against a rubric; score which of twenty rewrites a user most likely meant; gate a pipeline on whether a document is about X; give an ordinal quality signal on an agent's last step. It answers in tens of milliseconds of proxy-reported service time, under a strict output contract that held in all but 19 of ~16k benchmark requests, with probabilities attached, and it costs $0.28 per ten-thousand-odd hard questions - numbers a reasoning model cannot approach because it is doing a different job.

Our own favorite evidence that this is a real thing and not a demo: the model that cannot hallucinate still confidently picks wrong answers about a quarter of the time on physics, still wants you to believe it came from OpenAI, and turned out to be a tokenizer fingerprint nobody could match to an existing model. It earns attention by being useful, not by being frontier. That is a perfectly good thing to be. Just say so.

Notes

  1. Jev's responses report an “output” token count in usage - the serialized response JSON, counted but never billed; every output token is priced at zero. Every Jev dollar figure in this report is measured from billing usage in the published run artifacts, not estimated. [back]
  2. InstructGPT, of which Almeida was a coauthor, is best known for applying RLHF to GPT-3 and establishing a standard three-stage training pipeline, subsequently used in ChatGPT. However, RLHF did not originate with InstructGPT. It was first formulated by Russell and Ng (1998-2000), demonstrated through TAMER (2008-2012), PbRL (2011-2012), advanced with the 2017 seminal “Deep Reinforcement Learning from Human Preferences” (Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei), transitioning to natural language in Zeigler et al (2019) and Stiennon et al (2020), before being applied to GPT-3 with InstructGPT. GPT-3 in turn was enabled by numerous technologies: dense vector representations (Word2Vec, GloVe), subword tokenization / BPE, contextualized embeddings (ELMo), recurrent neural networks themselves, LSTMs and GRUs to solve the vanishing gradient problem, seq2seq, additive and scaled attention, the seminal work on Transformers (2017), autoregressive decoder models, self-supervised pretraining, transfer learning, empirical scaling laws, in-context learning / few-shot prompting, SFT, mixed-precision training, distributed parallelism frameworks (Megatron-LM, pipeline parallelism, ZeRO / DeepSpeed), and many more. “I co-invented” implies one of a small subset; ChatGPT has thousands of fathers. [back]
  3. No attempt was made to inspect, extract, or reconstruct the model itself - no weights, no gradients, no serving internals. Everything here comes from what the public API returns: its answers, its probability vectors, its token counts, its latency headers, and its billing usage. [back]
  4. Every probability returned across all runs landed exactly on a 0.01 grid - thousands of values, zero off-grid. That is a fixed-precision read-out, not raw token-level logits; it also means any reasoning built on tiny probability differences is reading noise. [back]
  5. Full lattice forensics: scripts/report/lattice_forensics.py -> data_report/lattice_forensics.json, over 704,277 values in 7,887 published vectors at K=2..255 (probe rows, phase-1 raws, and the Talk program's per-rotation large-menu distributions). Zero off-grid values; displayed sums are only ever 0.99 or 1.00 and never exceed 1.00, against the +-0.04 scatter independent per-value rounding would give at K=255 - so the display pipeline applies a bounded one-sided correction (or apportions integer hundredths), leaving a 0.01 shortfall on ~46% of flat large-K vectors. Every two-option vector sums to exactly 1.000 (complement emission). The grid is a rounding step, not a hard floor: probabilities below 0.005 display as exactly 0.00 where one answer is confident (common), while a flat vector keeps a 0.01 floor standard. The live identical-option probe (runs_archprobe/probe2/) confirmed sums of exactly 1.000 at every option count up to 12 and 0.99-1.00 at 255; the exact apportionment rule at the boundaries remains open. [back]
  6. The probability surface is post-processed, not raw. Every value sits on the 0.01 grid; nine published vectors return a choice that is not the argmax of their own displayed table, always at exactly one quantum - round-to-nearest cannot reorder options, so the decision is computed at pre-display precision; and across 2,289 published choice vectors the confidence field equals the chance-corrected top probability (p_max - 1/K) / (1 - 1/K) to within two quanta (1,405 exactly, none beyond). Score answers use a different shape statistic; noul answers carry no confidence at all. Re-derived from the published raw responses by scripts/report/arch_audits.py -> data_report/arch_audits.json. [back]
  7. Ordering numbers: runs_live/token_talk_orderprobe.json (11 orderings x 6 repeats of a frozen 254-option step) and docs/token-talk-findings.md sections 11-12; the rotation-ensemble result (K>=6 cancels the bias, Spearman 0.91-1.00 against the 12-rotation reference) is section 12. The repeat-noise band (TVD 0.03-0.12) is measured in the same file. The 419-item shuffle audit is runs_benchmark/bench-option_rotations-* paired against the native run (data_report/arch_audits.json). [back]
  8. External scores use their publishers' protocols, which are not Jev's: Vals MMLU-Pro is 5-shot with chain-of-thought; Artificial Analysis GPQA runs with reasoning enabled; ARC-AGI-2 rows are grid-production pass@2 on the semi-private set; Scale HLE rows cover the full text-only subset, partly with tools. Jev answered directly, with no reasoning and no tools. The charts show position, not a matched race. [back]
  9. MATH-500 and ARC-AGI-2 are protocol conversions, not native runs: Jev cannot generate a worked solution or a completed grid, so those problems were re-encoded as multiple-choice and per-cell questions. Those numbers say what the model knows, not what it can produce. The digit read-out variant is adapter-specific: its rubric is positional and conjunctive (every digit of the decimal must be picked in order), so the MCQ-versus-digit gap measures that adapter's recognition-to-production burden, not an isolated generation tax. The paired evidence is the 224 MATH-500 items present in both recorded runs: 0.817 as MCQ versus 0.103 digit-by-digit (19 both right, 164 MCQ-only, 4 digit-only, 37 neither; exact McNemar p below 1e-30), data_report/benchmark_diagnostics/paired_diagnostics.json. Baseline format failures are a separate axis from knowledge and are counted separately everywhere (strict-format failures recovered versus unrecovered). [back]
  10. Two scorings of the same responses, kept descriptive. Greedy takes the model's argmax option. The probability-weighted column is the average probability Jev itself put on the gold option: that is the model-implied expected success of stochastically selecting by its displayed probabilities, not a predicted argmax accuracy and not a calibration statement (the reliability table below tests calibration directly). Greedy exceeding weighted nearly everywhere shows only where the displayed mass sits relative to the chosen answer. Nothing here warrants an underconfidence or “over-dispersion” reading, and on HLE the calibration evidence runs the other way (mean top displayed probability ~0.634 against 0.219 accuracy). [back]
  11. Calibration numbers come from the repaired offline diagnostics (data_report/benchmark_diagnostics/calibration.json; deterministic, bootstrap seed 20260930). The reliability bins use the SELECTED answer's displayed probability, whose mean is the model-implied expected accuracy of taking that choice, and compare it to observed accuracy per bin. Displayed probabilities are rounded to two decimals and vector sums are only ever 0.99 or 1.00, so bin boundaries carry a ±0.005 caveat. Calibration is dataset-dependent and the high-confidence tail is where it hurts: on HLE the mean top displayed probability is ~0.634 against 0.219 accuracy (overconfident), while MMLU-Pro reads near-calibrated. Prefer task-specific calibration over a global conservative threshold. [back]
  12. The size-estimate machinery lives in data_report/size_estimate.json (scripts/report/size_estimate.py): both angles, the full assumption grid (MFU 0.25-0.45; the late-2026 fleet's BF16 nameplate ~400-2250 TFLOPS per device times a quantization multiplier of 2-4x (fp8/int4-FP4) = ~800-9000 effective; 1-4 shards; FLOPs = 2*N_active per token, attention adding ~15% or less at these lengths; single-stream conservative), the reconciling readings (quantized dense / MoE / distilled small dense), and what would tighten each. Two earlier passes were wrong, both labeled in the file: one assumed bf16 on A100-class hardware (~2-4x too low); the second treated the fleet's BF16 numbers as already effective (~2-4x too low again) - the audit caught it. The MFU/peak ranges are industry-standard serving assumptions, not repo measurements. All size figures here are hardware-dependent scenario calculations: service-time slopes are not hardware throughput measurements, so every parameter band is an illustration conditional on assumed hardware, quantization and utilization, not a reliable parameter estimate. These illustrative calculations use TypeSafe's reported input-token counts, which we have no evidence to dispute. [back]
  13. Each negation is a measurement, not a vibe. Compute grows linearly with prompt tokens - a cache would not prefill - and the knowledge horizon on the sampled dated-fact probes reads 2024-era strong and 2025 weak, which a live retriever would not. Repeats of identical payloads differ (TVD 0.03-0.12), which a lookup would not. The proxy-reported upstream service time stays 73-86 ms from concurrency 1-32, and four batched questions answer in ~261 ms vs ~1,117 ms separately - timings that disfavor a conventional frontier-API wrapper; proximity to the provider's own services, internal routing or custom billing remain possible, and a tokenizer mismatch cannot rule out reused weights. The usage counter follows a tokenizer matching none of 173 open signatures, including both OpenAI encodings. And $0.042/M with free output sits one to two orders of magnitude below every 2026 flagship's input list price (roughly 30-240x, before counting output they bill at 4-5x their input); reselling conventional frontier inference on those terms is disfavored. Paths: runs_archprobe/, runs_live/FINDINGS.md, data_report/costs.json. [back]
  14. The corrected ARC-AGI-2 choice/score runs (runs_benchmark_ext_rerun1) record the diagnostic protocol explicitly: oracle output shape and palette are an ORACLE_DIAGNOSTIC_PROTOCOL choice, not the official evaluation, and each cell question identifies its test grid. Per-cell numbers are adapter diagnostics; the official task criterion (0 of 120 tasks, all cells of all test grids correct) is scored on its own run and is assisted adaptation rather than an official matched benchmark. [back]
  15. Baseline revision status (v4r1): the ACTIVE summary (data_report/baselines/v4r1/public_summary.json) supersedes the older v2/v3 baseline summaries. 59 of 59 (model, dataset) cells are complete over their full sampled denominators (31,038 requested rows overall) and are the only cells used in capability comparisons. The status is versioned (v4r1): legacy v2/v3 recovered scores remain labeled as legacy and are never presented as current. Costs: $25.70 provider-reported original rows + $24.24 settled replacements ($1.573018 held) + $1.16 Typesafe ARC usage-at-tariff (not invoice-verified), settled total $25.397157 against the $80 combined cap; unknown billed costs are held as unknown, never treated as free. [back]
  16. Re-rendering this page (scripts/report/build_report.py over the published aggregates) re-derives prose and figures from those aggregates byte-for-byte; it is not an independent re-scoring. Independently re-scoring would require the private raw responses and the licensed item sets, which are not republished; the freeze records carry SHA-256 hashes that pin their identity exactly. Per-item prediction evidence (item ids, correctness, displayed probabilities) is held in the working tree under runs_matched_cheap/v4r1/active_items.jsonl and the stage per_item lists; it is excluded from the published bundle under the same item-id policy that strips per_item lists at publish time. [back]
  17. A robustness audit: MMLU-Pro items re-run with option order rotated, to check the headline numbers are not an artifact of where the answer sits in the list. The design is 140 unique base items with 3 variants each (native plus rotations; 419 scored observations of 420 requested). The observations are clustered on their base items and are not independent. The cluster-aware analysis (per-base differences, 10,000 cluster bootstraps, seed 20260930) gives mean native-minus-rotated accuracy +0.36 points with 95% CI [-2.50, +3.33] points: consistent with no systematic order effect on these items at this resolution, and explicitly not an equivalence proof. The naive independent McNemar p = 0.83 in older artifacts assumes independence the design does not have; it is not quoted here as evidence. [back]
  18. Matched cheap-model baselines: scripts/benchmark/run_matched3.py (v3, spec-compliant) over run_cheap_matched2.py (v2); artifacts in runs_matched_cheap/ (v3/ preferred, v2/ kept for provenance). Identical frozen items to the Jev runs (MMLU-Pro 1,000-item stratified subset, seed 20260920; GPQA 196; MATH-500 choice 261; HLE MC 494; ARC-Challenge 1,172 for several models), one attempt per item, transport-level backoff only, provider-reported costs. v3 wire policy: NO temperature, top_p, seed or reasoning parameters - each model runs at its documented provider defaults, per the audit's out-of-spec warning; max_tokens is a bill guard, not a thinking budget. (v2 pinned temperature 0 + effort low, which is out of spec for reasoning models and depressed their scores and costs; v2-protocol carryover cells, kept only where v3 never re-ran them, are listed per cell in data_report/baselines/v4r1/active_source_map.json (visible carryovers, not just Qwen3.8 Max MMLU) and are labeled wherever they appear.) Answers are parsed strictly and, failing that, recovered by a documented deterministic parser applied identically to every model; unrecovered answers count as wrong. The cost gap versus vals.ai rows for the same model is protocol, not metering: vals runs 5-shot chain-of-thought at max effort (2-50x the tokens), ours are one-shot at defaults - both figures are what the providers actually billed. Budget $80 total authorized by the project owner; spend in the live_summary files. Era colors use curated/estimated release dates, labeled in tooltips. [back]
  19. External costs are the vals.ai platform's measured cost per test - the same harness that measured the scores, chain-of-thought tokens included, so reasoning counts against the models that use it. Jev's costs are measured from our billing (reported input tokens x $0.042/M; output free). The earlier draft's list-price estimate method and its arithmetic remain in data_report/costs.json for provenance; the charts no longer use it. [back]