Huddleston Personal ComputerResearch / Decision Aux

BEARHUDDLESTON.DEV / DECISION AUX

Decision Aux:
consolidated PoC findings.

Jev made decisions faster, but adding it did not demonstrate better agent work. The useful task-level improvement came from supplying source files without Jev.

Decision Aux is an optional helper step: ask a model which model should do the task, which files it needs, or which skills to load. Jev is one candidate for that helper; Astra does the main work in the source and skill trials.

Just want the recommendation? Read our integration assessment →

Keep in mind: these are small tests with made-up tasks, not proof of everyday performance. Each pilot (PoC) asks a different question. Four are complete; only the combined test, PoC 3, is on hold and was never run. A quicker helper call is not necessarily a quicker task.

Model choice · completed

Did Jev choose the faster task model?

Result. Jev moved 0/16 tasks to Luna, the faster model selected in an earlier screen.

Meaning. It kept Astra every time and added a routing call. This did not demonstrate faster work through model selection.

Data and method: model choice

01 / CHOOSE THE GENERATOR

PoC 1 / COMPLETE · Pilot completed.

Jev kept Astra on every primary task.

The router ran once, before constructing a fresh agent. It could choose the screened faster generator, Luna, or keep Astra. In the Jev arm, every task still ran on Astra.

Astra only
19.31s
Luna only
16.96s
Rule router
18.24s
Astra router
19.57s
Jev router
18.80s

Seconds per task; lower is faster. Mean shown initially. Each arm has 16 attempts. Times include routing, agent construction and the ordinary tool loop.

ArmArtifactsMean secondsMedian secondsMain requestsAux requests
Astra only16/1619.31219.4911030
Luna only15/1616.96314.682980
Rule router16/1618.23517.564990
Astra router16/1619.56920.5559616
Jev router16/1618.79519.8719916
Astra only0/16 Luna
Luna only16/16 Luna
Rule router0/16 Luna
Astra router2/16 Luna
Jev router0/16 Luna

Each square is one task. Green = Astra; amber = Luna. A router cannot demonstrate faster-generator allocation when it never leaves Astra.

Interpretation: Jev's raw mean was lower than Astra-only, but its median was higher. It added about 0.43 seconds of routing work without allocating any task to Luna. That is not demonstrated generation-routing acceleration.

Quality miss, fallback and uncertainty
  • Luna-only passed 15/16. One service-order artifact was incorrect; the model's self-check accepted the same incorrect order. Four model tool errors were retained with model-driven recovery.
  • The Astra router timed out twice. Jev had two combined abstained/invalid decisions. All four fell back to Astra; the saved reason does not distinguish Jev abstention from invalidity.
  • The rule router also kept Astra in all 16 cases. Its lower mean cannot be credited to faster-generator allocation either.
  • Jev saw the request, not the task's specification or source. Cautious policy and insufficient task evidence are possible explanations, not established causes.
Treatment minus controlMean paired difference95% task-cluster interval
Luna only − Astra only-2.349s-6.272 to +2.911s
Rule router − Astra only-1.076s-1.917 to -0.235s
Astra router − Astra only+0.258s-2.060 to +2.252s
Jev router − Astra only-0.516s-2.095 to +0.809s

Negative differences favor the named treatment. Eight task clusters, two repeats each; 5,000 clustered bootstrap resamples. Intervals describe task variation here, not every provider/cache effect.

Generator calibration — separate from primary evidence

Three calibration tasks × two repeats × four generators. Eligibility required every attempt to pass, the requested model identity, and real tool execution; lowest mean among eligible models selected Luna. No primary task was used to tune the choice.

Model via APIPassesMean secondsEligible
@chatgpt/gpt-6-astra6/615.747Yes
@chatgpt/gpt-5.4-mini5/615.947No
@chatgpt/gpt-5.6-luna6/610.067Yes
@claude/claude-haiku-4-5-202510016/619.570Yes
Every recorded primary attempt
EpisodeTaskRepeatArmArtifactSecondsMainAux
e046p011Rule routerPASS12.03440
e076p011Luna onlyPASS17.34750
e042p011Jev routerPASS12.93141
e041p011Astra routerPASS11.38241
e061p011Astra onlyPASS10.48140
e040p012Rule routerPASS10.86240
e062p012Luna onlyPASS8.89050
e060p012Jev routerPASS11.66041
e033p012Astra routerPASS10.50541
e032p012Astra onlyPASS15.72160
e073p021Rule routerPASS12.49750
e057p021Luna onlyPASS10.24150
e069p021Jev routerPASS12.20251
e045p021Astra routerPASS16.90851
e025p021Astra onlyPASS11.79850
e056p022Rule routerPASS14.87250
e036p022Luna onlyPASS8.80050
e053p022Jev routerPASS14.87751
e031p022Astra routerPASS18.33051
e037p022Astra onlyPASS14.02850
e019p031Rule routerPASS14.70450
e077p031Luna onlyPASS13.08560
e038p031Jev routerPASS15.80051
e043p031Astra routerPASS14.36251
e075p031Astra onlyPASS22.69470
e026p032Rule routerPASS21.44770
e058p032Luna onlyPASS16.04570
e021p032Jev routerPASS20.87871
e071p032Astra routerPASS15.06851
e008p032Astra onlyPASS19.53670
e013p041Rule routerPASS15.96060
e022p041Luna onlyFAIL14.98450
e047p041Jev routerPASS18.12361
e064p041Astra routerPASS23.22281
e015p041Astra onlyPASS16.23160
e063p042Rule routerPASS17.85560
e055p042Luna onlyPASS15.20770
e059p042Jev routerPASS16.85561
e068p042Astra routerPASS17.99261
e079p042Astra onlyPASS19.00960
e035p051Rule routerPASS23.85870
e017p051Luna onlyPASS14.38070
e012p051Jev routerPASS27.98671
e067p051Astra routerPASS22.29761
e072p051Astra onlyPASS22.26770
e003p052Rule routerPASS19.97060
e009p052Luna onlyPASS24.50180
e078p052Jev routerPASS20.29561
e066p052Astra routerPASS23.51261
e052p052Astra onlyPASS24.27660
e051p061Rule routerPASS17.27370
e001p061Luna onlyPASS15.84460
e010p061Jev routerPASS20.02171
e018p061Astra routerPASS20.10971
e024p061Astra onlyPASS19.86570
e011p062Rule routerPASS17.12570
e006p062Luna onlyPASS10.07950
e005p062Jev routerPASS20.66671
e014p062Astra routerPASS21.00271
e002p062Astra onlyPASS19.44760
e039p071Rule routerPASS30.36270
e080p071Luna onlyPASS41.10280
e027p071Jev routerPASS21.85661
e029p071Astra routerPASS30.79061
e020p071Astra onlyPASS26.62170
e023p072Rule routerPASS19.90760
e049p072Luna onlyPASS38.34970
e054p072Jev routerPASS19.72061
e007p072Astra routerPASS23.50061
e048p072Astra onlyPASS24.69080
e044p081Rule routerPASS19.59280
e004p081Luna onlyPASS10.86660
e050p081Jev routerPASS23.80891
e034p081Astra routerPASS22.22381
e070p081Astra onlyPASS19.21580
e016p082Rule routerPASS23.44290
e028p082Luna onlyPASS11.68860
e065p082Jev routerPASS23.04991
e074p082Astra routerPASS21.90781
e030p082Astra onlyPASS23.10580

PoCs 1–2 / the routing and retrieval mechanisms

Two separate experiments: choose the generator before a task, or supply exact source before Astra.01 / ROUTEWhich generator?Astra or Luna02 / RETRIEVEWhich source files?Astra stays fixedFresh task requestRules / LLM / JevNormal agent loopSame source repositoryAll / lexical / JevNormal agent loop

01 / ROUTETask → choose Astra or Luna → normal agent loop.

02 / RETRIEVESource files → bundle evidence → normal Astra loop.

Source selection · completed

Did Jev help find the right source files?

Result. Giving Astra all the source files, without Jev, cut mean task time by 14.4%. All files fit within the budget.

Meaning. Simple bundling helped here. Jev’s file ranking did not demonstrate an extra benefit over the simpler selection methods.

Data and method: source selection

02 / SUPPLY THE EVIDENCE

PoC 2 / COMPLETE · Pilot completed.

All-source bundling was the useful control.

Astra stayed fixed. Every task had 20 short source files, all within the 9,000-character source budget. One arm supplied all files; two shortlisted up to six using the same import expansion. Ordinary Astra kept unrestricted search, reads and batching.

Ordinary Astra
17.26s
All-source bundle
14.77s
Lexical bundle
16.56s
Jev bundle
17.10s

Seconds per task; lower is faster. Mean shown initially. Ranking, source reads and agent initialization are included.

ArmArtifactsMean secondsMedian secondsMain requestsAux requests
Ordinary Astra16/1617.26217.8511100
All-source bundle16/1614.76815.355880
Lexical bundle16/1616.56017.0471000
Jev bundle16/1617.10316.66910716

14.4% lower mean time

All-source bundling: 14.77 seconds versus ordinary Astra's 17.26 seconds. Every artifact passed. Main requests fell from 110 to 88: a 20% reduction.

No incremental Jev win

Jev averaged 17.10 seconds versus lexical selection's 16.56 seconds. It used 107 main requests plus 16 ranking requests. Its mean retrieval overhead was about 0.51 seconds per task.

Treatment minus controlMean paired difference95% task-cluster interval
All-source bundle − Ordinary Astra-2.494s-4.285 to -0.743s
Lexical bundle − Ordinary Astra-0.702s-2.381 to +1.040s
Jev bundle − Ordinary Astra-0.159s-1.768 to +1.249s
Jev bundle − Lexical bundle+0.543s-1.100 to +2.239s
Jev bundle − All-source bundle+2.335s+1.306 to +3.640s

Task-cluster bootstrap intervals; negative means faster. Jev–lexical spans both benefit and harm. Every task-family mean favored all-source bundling over Jev in this run. This does not establish general quality equivalence.

Fewer reads did not guarantee fewer rounds.

ArmRead callsSearch callsReads of bundled filesMean injected characters
Ordinary Astra1521600
All-source bundle871558040
Lexical bundle1361203190
Jev bundle1198623067

Jev supplied governing code that the lexical shortlist often missed. Astra nevertheless reread bundled files 62 times. Those reads can be legitimate verification; fewer search/read calls alone did not produce fewer inference requests.

Same task, different discovery work

All-source: 3 main requests, 8.81s. Jev: 7 main requests, 18.30s. This single example illustrates the mechanism, not a separate aggregate claim.

Method limits and task-family differences
  • The source trees were small: all 20 files fit. This is not a large-repository RAG benchmark.
  • Lexical selection used BM25 over the full task request. Shared instruction wording dominated many selections, favoring README and operational prose. It is not evidence against a tuned, task-focused retrieval baseline. No after-the-fact retuning was performed.
  • All-source bundles averaged 4,800 source characters. Framing increased the injected text to about 8,040 characters. The six-file Jev bundle averaged about 3,067 injected characters.
  • Jev ranked successfully in all 16 primary episodes. No provider, tool, interruption or fallback errors, repairs or reruns were recorded.
  • Exact source text was appended once to the initial user request. The system prefix, toolset and Astra identity were unchanged; original sources remained accessible.
Task familyAll-source − ordinaryLexical − ordinaryJev − ordinary
r01-0.63s+0.45s+1.25s
r02+0.70s+0.13s+1.84s
r03-5.45s-1.88s-4.12s
r04-3.85s-2.92s-0.36s
r05-3.63s+0.08s-3.02s
r06-6.42s-5.15s-0.39s
r07-1.00s+3.71s+2.33s
r08+0.33s-0.03s+1.20s
Every recorded primary attempt
EpisodeTaskRepeatArmArtifactSecondsMainAux
e050r011All-source bundlePASS12.78750
e051r011Lexical bundlePASS15.64160
e028r011Jev bundlePASS14.19461
e056r011Ordinary AstraPASS15.06270
e063r012All-source bundlePASS12.18150
e036r012Lexical bundlePASS11.47850
e006r012Jev bundlePASS14.52061
e019r012Ordinary AstraPASS11.15740
e009r021All-source bundlePASS16.12360
e027r021Lexical bundlePASS15.24870
e064r021Jev bundlePASS14.93461
e011r021Ordinary AstraPASS15.96670
e054r022All-source bundlePASS14.93960
e025r022Lexical bundlePASS14.67260
e029r022Jev bundlePASS18.41371
e003r022Ordinary AstraPASS13.69450
e026r031All-source bundlePASS16.09760
e021r031Lexical bundlePASS20.10380
e001r031Jev bundlePASS15.90061
e032r031Ordinary AstraPASS20.61180
e035r032All-source bundlePASS15.38050
e005r032Lexical bundlePASS18.51070
e014r032Jev bundlePASS18.24571
e010r032Ordinary AstraPASS21.76780
e022r041All-source bundlePASS14.77350
e047r041Lexical bundlePASS21.53980
e020r041Jev bundlePASS21.39671
e058r041Ordinary AstraPASS20.17080
e053r042All-source bundlePASS15.33160
e030r042Lexical bundlePASS10.42340
e049r042Jev bundlePASS15.68871
e037r042Ordinary AstraPASS17.64180
e052r051All-source bundlePASS15.27660
e015r051Lexical bundlePASS21.09480
e031r051Jev bundlePASS16.84971
e008r051Ordinary AstraPASS18.98880
e007r052All-source bundlePASS16.84660
e046r052Lexical bundlePASS18.45370
e062r052Jev bundlePASS16.48861
e048r052Ordinary AstraPASS20.39280
e024r061All-source bundlePASS8.81130
e033r061Lexical bundlePASS12.81340
e055r061Jev bundlePASS18.30571
e061r061Ordinary AstraPASS16.83060
e012r062All-source bundlePASS13.24050
e038r062Lexical bundlePASS11.76740
e016r062Jev bundlePASS15.81071
e034r062Ordinary AstraPASS18.06170
e057r071All-source bundlePASS15.39260
e018r071Lexical bundlePASS22.48880
e042r071Jev bundlePASS17.16171
e040r071Ordinary AstraPASS20.38480
e041r072All-source bundlePASS16.71860
e045r072Lexical bundlePASS19.05470
e059r072Jev bundlePASS21.61981
e004r072Ordinary AstraPASS13.73350
e017r081All-source bundlePASS16.76460
e002r081Lexical bundlePASS13.00740
e013r081Jev bundlePASS18.26971
e060r081Ordinary AstraPASS13.28760
e039r082All-source bundlePASS15.63260
e023r082Lexical bundlePASS18.67170
e044r082Jev bundlePASS15.86261
e043r082Ordinary AstraPASS18.45070

Combined test · on hold

Did model choice and source selection help together?

Result. This experiment was never run.

Meaning. The separate studies cannot establish a combined speedup. There is no result to interpret.

Why there is no combined result

03 / COMBINE ROUTING AND RETRIEVAL

PoC 3 / ON HOLD · Never run.

No combined result.

The neither/routing/retrieval/both comparison has not been run on the reserved common holdout. PoCs 1 and 2 cannot be combined after the fact to claim synergy or a combined speedup. This experiment remains paused.

Decision speed · completed

Could Jev answer the helper call faster?

Result. Jev returned the same correct labels as Astra with 72.3% less median call time in this test.

Meaning. That is a real call-level latency win, not evidence of better answers or faster completed tasks. The runs used different APIs at different times.

Data and method: decision speed

04 / JEV VS ASTRA BACKEND

PoC 4 / COMPLETE · Pilot completed.

Jev returned the same labels sooner.

The backend comparison used 72 unique, clear synthetic cases. Both Jev and Astra returned all 72/72 labels correctly. Each backend made 102 calls including CLI, warmup, repeat and padding controls; those controls are not additional independent accuracy samples.

Separate sequential runs at different times, different APIs and client behavior, and provider-default sampling for Astra limit the comparison. Timings include network and validation, not just model compute. No verified dollar comparison is available.

Median decision-call latency

Jev375 ms
Astra1350 ms

Shared scale: 0–3,000 ms. The source report records 72.3% less call time for Jev, not a task-completion speedup. Primary latency uses the 72 unique cases.

Call-level win, not a quality win. The corpus saturated both backends. It establishes neither a quality advantage nor an end-to-end agent benefit. It does show that Jev can make these bounded decision calls faster.

Jev 1.13.0 versus your configured Astra model via API. This compares the two Decision Aux backend paths—not autonomous-agent performance with the feature switched on and off.

Same 72 casesHermes Decision AuxJev native APIorAstra LLM APISame validation

What this supports: Jev reduced latency on these bounded judgments. What it does not support: a quality advantage, production accuracy claim, or end-task improvement.

Measured comparison
MetricJevNo Jev: Astra
Valid responses, final run102 / 102102 / 102
Unique labels correct72 / 7272 / 72
Median runtime call375 ms1350 ms
P95 runtime call464 ms1606 ms
Maximum runtime call534 ms2667 ms
Fresh CLI process range · n=3732–856 ms2184–3378 ms
Repeat label agreement12 / 1212 / 12
Short-padding label agreement12 / 1212 / 12
Input / prompt tokens · all 10256,42952,201
Output / completion tokens · all 1022,7882,424
Estimated cost$0.00237Not verified

Primary latency uses only the 72 unique cases. The other 30 calls are CLI/warmup/repeat/padding controls, not additional independent accuracy samples. Different tokenizers and response protocols: token counts are not a cost ratio.

The latency distributions
Jev and Astra latency on the same 72 unique inputs, a shared linear millisecond scale0 ms500 ms1000 ms1500 ms2000 ms2500 ms3000 msJevevidence-03: 323.5 msevidence-17: 324.4 msskill-18: 330.2 msevidence-12: 334.1 msseverity-08: 334.1 msskill-09: 336.8 msevidence-05: 338.4 msseverity-17: 339.6 msskill-24: 341.8 msseverity-04: 343.2 msskill-01: 343.8 msevidence-07: 344.2 msevidence-11: 344.9 msevidence-02: 345.3 msskill-02: 349.9 msseverity-07: 352.2 msskill-08: 354.2 msskill-20: 354.6 msseverity-16: 355.0 msseverity-19: 356.1 msevidence-18: 356.3 msseverity-20: 357.6 msseverity-12: 358.3 msskill-11: 358.9 msevidence-19: 362.2 msseverity-06: 366.1 msskill-17: 366.3 msskill-16: 367.4 msevidence-09: 367.5 msseverity-21: 370.9 msseverity-23: 371.6 msskill-21: 371.7 msskill-07: 371.9 msevidence-10: 374.3 msskill-19: 374.4 msskill-15: 374.6 msevidence-16: 374.7 msskill-10: 375.6 msseverity-18: 376.2 msskill-13: 378.1 msevidence-15: 383.6 msseverity-15: 383.6 msevidence-24: 385.3 msevidence-04: 387.4 msskill-23: 387.4 msevidence-23: 388.0 msseverity-24: 388.7 msevidence-01: 389.5 msevidence-20: 390.0 msskill-03: 391.3 msevidence-08: 394.6 msevidence-21: 395.0 msseverity-10: 396.3 msskill-12: 397.0 msevidence-13: 398.8 msevidence-14: 399.6 msseverity-05: 406.0 msseverity-02: 406.9 msskill-05: 409.4 msskill-22: 416.6 msseverity-03: 418.8 msskill-14: 419.4 msevidence-22: 433.5 msseverity-11: 436.1 msevidence-06: 437.2 msseverity-14: 449.4 msseverity-01: 450.3 msseverity-22: 460.9 msskill-06: 467.3 msseverity-09: 483.9 msseverity-13: 484.0 msskill-04: 534.1 msmedian 375Astraskill-02: 1160.0 msskill-15: 1164.8 msskill-14: 1210.8 msevidence-16: 1219.0 msskill-01: 1222.2 msseverity-07: 1227.2 msevidence-23: 1240.0 msseverity-13: 1247.1 msseverity-08: 1247.8 msskill-06: 1267.6 msevidence-01: 1269.0 msskill-08: 1269.3 msskill-07: 1271.2 msseverity-16: 1278.1 msskill-11: 1282.2 msevidence-10: 1284.9 msseverity-06: 1288.8 msskill-21: 1289.0 msevidence-07: 1291.5 msskill-20: 1291.6 msevidence-14: 1294.6 msevidence-08: 1294.8 msskill-19: 1295.2 msseverity-20: 1299.8 msevidence-13: 1301.0 msevidence-22: 1307.1 msskill-17: 1308.6 msevidence-06: 1315.3 msevidence-21: 1316.1 msseverity-12: 1317.4 msskill-18: 1329.6 msevidence-03: 1340.3 msseverity-24: 1340.6 msskill-12: 1345.4 msseverity-02: 1346.3 msskill-16: 1348.6 msseverity-03: 1352.4 msseverity-05: 1352.8 msskill-22: 1356.4 msevidence-05: 1366.5 msevidence-09: 1370.8 msseverity-10: 1391.7 msseverity-17: 1395.4 msevidence-04: 1396.0 msseverity-04: 1397.9 msskill-10: 1401.0 msevidence-12: 1402.3 msseverity-09: 1404.6 msseverity-19: 1405.6 msevidence-19: 1412.3 msskill-13: 1414.4 msseverity-22: 1423.9 msskill-03: 1423.9 msskill-09: 1425.9 msevidence-02: 1427.7 msevidence-20: 1439.1 msevidence-15: 1442.5 msseverity-18: 1442.8 msskill-24: 1445.1 msskill-05: 1445.2 msevidence-17: 1445.5 msevidence-11: 1459.3 msseverity-23: 1460.0 msseverity-14: 1470.4 msseverity-15: 1506.3 msseverity-01: 1510.4 msevidence-18: 1570.6 msseverity-21: 1593.5 msskill-04: 1620.3 msseverity-11: 1982.7 msskill-23: 1984.8 msevidence-24: 2667.4 msmedian 1350

Astra's median was 3.60× Jev's. Jev was faster on all 72 paired inputs; median paired difference 955 ms. These were separate sequential runs, so time, service load and routing can confound the comparison.

Quality: a tie on every task
Skill selectionJev 24/24Astra 24/24
Evidence checksJev 24/24Astra 24/24
SeverityJev 24/24Astra 24/24

Both had zero false acceptances among 12 unsupported-evidence examples, and zero severity errors among 24 examples. The corpus is too easy to distinguish judgment quality. No probabilities or confidence were fabricated for the LLM backend.

The comparison found a real compatibility bug

The initial Astra smoke failed: the relay returned HTTP 400, upstream_unavailable, “Model unavailable.” A plain request worked. Single-variable probes showed that temperature: 0 caused the rejection; strict JSON schema without explicit temperature worked.

Fixed the new LLM backend to request the provider's default sampling setting (temperature=None). Kept strict schema and local validation. Expected labels, task instructions, Jev code, timeout and schedule were unchanged. Two regression checks failed before the fix; afterwards 874 tests passed across 58 selected files.

The failed smoke remains in llm-v2-001. Nine diagnostic inference attempts, including five successes, are excluded from benchmark scores and timings; their known usage and failures are included in the evidence. This is not a claim that the first unmodified baseline run succeeded.

Fairness, timing and limits
  • Backend E2E comparison, not full autonomous-agent on/off A/B.
  • Clear synthetic corpus saturates both models: no quality advantage established.
  • Same corpus and schedule; separate runs at different times, not interleaved.
  • Jev request-scoped HTTP clients versus host-cached LLM client; network/routes differ.
  • Labels and task rubrics identical; each adapter uses its native request/response protocol.
  • Astra uses provider-default sampling after compatibility fix; schema and local validation remain enabled.
  • Only 72 distinct cases per backend; repeated phases do not increase sample independence.
  • Five-second timeout and concurrency one; no claim about load behavior or production tails.
  • No verified dollar comparison without API pricing.
  • TypeSafe MCA restricts benchmark publication.

The ordinary LLM backend uses Hermes's cached client; the Jev adapter builds a request-scoped client. Each native adapter translates the same bounded questions into its own protocol. Timings include network and validation, not just GPU/model time. Initial plugin discovery is excluded from runtime timings but included in fresh CLI wall time.

Observed all 99 in-process baseline calls: one HTTP response per call, no recovery-log events. The three CLI calls exited successfully but their internal request count was not instrumented. Backend credentials were isolated; the no-Jev profile contained no Jev key or plugin.

Model: @chatgpt/gpt-6-astra via API. This is the configured main-model baseline, not a small optimized classification model. No API price was established, so no dollar-savings ratio is reported.

All phase timing populations
PhaseNJev median msAstra median msJev range msAstra range ms
Fresh CLI processes3 each751.342354.29732–8562184–3378
Initial in-process calls3 each408.941448.10372–4431275–2100
Unique cases72 each374.651350.49324–5341160–2667
Repeat controls12 each368.641322.89354–4981210–2011
Short padding controls12 each366.541337.14331–5231197–1616

Small CLI/warmup samples do not establish reliable tail latency. Padding is roughly 1.5k tokens, not a broad long-context test.

Inspect all 72 paired labels and latencies
CaseExpectedJevAstraJev msAstra ms
evidence-01TrueTrueTrue389.51269.0
evidence-02TrueTrueTrue345.31427.7
evidence-03TrueTrueTrue323.51340.3
evidence-04TrueTrueTrue387.41396.0
evidence-05TrueTrueTrue338.41366.5
evidence-06TrueTrueTrue437.21315.3
evidence-07TrueTrueTrue344.21291.5
evidence-08TrueTrueTrue394.61294.8
evidence-09TrueTrueTrue367.51370.8
evidence-10TrueTrueTrue374.31284.9
evidence-11TrueTrueTrue344.91459.3
evidence-12TrueTrueTrue334.11402.3
evidence-13FalseFalseFalse398.81301.0
evidence-14FalseFalseFalse399.61294.6
evidence-15FalseFalseFalse383.61442.5
evidence-16FalseFalseFalse374.71219.0
evidence-17FalseFalseFalse324.41445.5
evidence-18FalseFalseFalse356.31570.6
evidence-19FalseFalseFalse362.21412.3
evidence-20FalseFalseFalse390.01439.1
evidence-21FalseFalseFalse395.01316.1
evidence-22FalseFalseFalse433.51307.1
evidence-23FalseFalseFalse388.01240.0
evidence-24FalseFalseFalse385.32667.4
severity-01Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact450.31510.4
severity-02Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact406.91346.3
severity-03Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact418.81352.4
severity-04Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact343.21397.9
severity-05Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact406.01352.8
severity-06Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact366.11288.8
severity-07Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact352.21227.2
severity-08Cosmetic only; no functional impactCosmetic only; no functional impactCosmetic only; no functional impact334.11247.8
severity-09Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists483.91404.6
severity-10Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists396.31391.7
severity-11Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists436.11982.7
severity-12Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists358.31317.4
severity-13Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists484.01247.1
severity-14Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists449.41470.4
severity-15Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists383.61506.3
severity-16Function degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround existsFunction degraded or unavailable, but a practical workaround exists355.01278.1
severity-17Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss339.61395.4
severity-18Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss376.21442.8
severity-19Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss356.11405.6
severity-20Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss357.61299.8
severity-21Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss370.91593.5
severity-22Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss460.91423.9
severity-23Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss371.61460.0
severity-24Core function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data lossCore function blocked without a workaround, or permanent data loss388.71340.6
skill-01githubgithubgithub343.81222.2
skill-02githubgithubgithub349.91160.0
skill-03githubgithubgithub391.31423.9
skill-04githubgithubgithub534.11620.3
skill-05githubgithubgithub409.41445.2
skill-06githubgithubgithub467.31267.6
skill-07pdfpdfpdf371.91271.2
skill-08pdfpdfpdf354.21269.3
skill-09pdfpdfpdf336.81425.9
skill-10pdfpdfpdf375.61401.0
skill-11pdfpdfpdf358.91282.2
skill-12pdfpdfpdf397.01345.4
skill-13xlsxxlsxxlsx378.11414.4
skill-14xlsxxlsxxlsx419.41210.8
skill-15xlsxxlsxxlsx374.61164.8
skill-16xlsxxlsxxlsx367.41348.6
skill-17xlsxxlsxxlsx366.31308.6
skill-18xlsxxlsxxlsx330.21329.6
skill-19nonenonenone374.41295.2
skill-20nonenonenone354.61291.6
skill-21nonenonenone371.71289.0
skill-22nonenonenone416.61356.4
skill-23nonenonenone387.41984.8
skill-24nonenonenone341.81445.1
Reproducibility and privacy

Corpus SHA-256: 2bdea01789121cf951d2b821202b342aca18d78154597f62d78b063105d7168f.
Jev run: live-v2-001.
No-Jev final run: llm-v2-002.

This report-only edition includes the comparison charts, methodology and paired-case table. Raw result files, manifests, the separate CSV, comparison script, diagnostic files, test output and source patch are not included as downloads.

Jev estimated cost uses reported input usage and TypeSafe's published $0.042/M input-token price, not invoice billing. TypeSafe MCA §2.3 restricts benchmark publication.

Measured backend E2E comparison. No fabricated model output. Normal Hermes installation and profile settings unchanged. Report checks cover DOM/layout, not pixel inspection.

Skill advice · completed

Did faster skill advice improve the finished work?

Result. Every group completed 8/8 trials, including the group with no helper call.

Meaning. Jev advised faster than Astra, but no task-quality gain was demonstrated. More skill loading was not an improvement by itself.

Data and method: skill advice

05 / AGENT SKILL ADVISOR

PoC 5 / COMPLETE · Pilot completed.

A faster advisor did not establish a better task outcome.

The full-agent pilot ran 24 trials: four tasks × two repeats × three arms, with Astra as the main model throughout. No decision call, Astra advice and Jev advice each had 8/8 artifact-complete trials. No task-quality gain was demonstrated.

Four synthetic tasks and a ten-skill catalog do not establish production benefit or large-catalog recall. Repeats add no task diversity. Recommendation and actual skill loading are separate. The no-decision-call baseline is essential: comparing only two advisors would miss whether advice helped at all.

Median full-agent task duration

No decision call20.57s
Astra advice24.77s
Jev advice21.17s

Shared scale: 0–40s. Includes the decision call, main model, tools and finalization; excludes initialization. These mixed-task medians are descriptive, not a causal speed estimate.

Advisor medians were 2870 ms for Astra and 395 ms for Jev. Total skill loads were 9 with no decision call, 11 with Astra advice and 16 with Jev advice. More loading was not a quality improvement.

Main model: @chatgpt/gpt-6-astra via API in every arm. Jev backend: jev-1.13.0.

One main model. The same four tasks. Three decision paths. Each trial ends in real files checked by deterministic code—not a model grading itself.

1 · Recommend or skipOff / Astra / Jev · same visible skill catalog
2 · Hermes does the workAstra main model · fresh isolated workspace
3 · Check artifactsCSV, JSON, Markdown, exact-byte copy
Keep the advisor off by default.

All arms completed every artifact correctly; this pilot did not demonstrate a task-quality gain from the extra call. The no-decision baseline is part of every future E2E comparison.

More skill loading is not necessarily better.

Jev advice: 16 loads · Astra advice: 11 · No decision call: 9. In the release task, Jev recommended an internal-changelog recipe and a Markdown-table recipe; the agent loaded those alongside the release exporter. The requested artifacts still passed, but advice added unnecessary workflow selection. This ten-skill catalog does not test large-catalog recall.

Outcomes, advisor work and tool usage

No decision call

8 / 8

artifact-complete trials

Task median
20.57s
Advisor median
Skill loads
9
Tool calls
61

Astra advice

8 / 8

artifact-complete trials

Task median
24.77s
Advisor median
2870 ms
Skill loads
11
Tool calls
63

Jev advice

8 / 8

artifact-complete trials

Task median
21.17s
Advisor median
395 ms
Skill loads
16
Tool calls
68
Complete task duration

Every dot is one trial. Eight per arm; four task types repeated twice. White ticks mark medians.

Full agent task duration in seconds, all eight trials per arm; shared zero-based scale0s10s20s30s40sNo decision calle008: t03-copy-note / 9.46se019: t03-copy-note / 10.23se024: t04-release-notes / 18.22se010: t04-release-notes / 18.86se006: t02-support-bundle / 22.29se017: t02-support-bundle / 25.54se001: t01-ledger-close / 25.75se015: t01-ledger-close / 27.41smedian 20.6sAstra advicee020: t03-copy-note / 11.15se009: t03-copy-note / 15.83se011: t04-release-notes / 19.58se022: t04-release-notes / 22.99se018: t02-support-bundle / 26.56se002: t01-ledger-close / 26.72se004: t02-support-bundle / 27.75se013: t01-ledger-close / 32.13smedian 24.8sJev advicee007: t03-copy-note / 12.41se021: t03-copy-note / 13.54se023: t04-release-notes / 18.19se012: t04-release-notes / 18.80se014: t01-ledger-close / 23.54se003: t01-ledger-close / 24.72se016: t02-support-bundle / 25.24se005: t02-support-bundle / 25.64smedian 21.2s

Includes decision call + main model + tools + finalization. Excludes initialization. Mixed task durations are descriptive, not a causal speed estimate.

What happened in each task?

Recommendation and actual skill loading are separate. An agent can ignore advice or load additional skills.

24 trials shown.

Task / repeatArmArtifactsTimeRecommendedActually loaded
ledger-closerepeat 1 · e003Jev advicePASS24.72scsv-money-close, csv-data-qualitycsv-money-close, csv-data-quality
ledger-closerepeat 1 · e002Astra advicePASS26.72scsv-money-close, csv-data-qualitycsv-money-close, csv-data-quality
ledger-closerepeat 1 · e001No decision callPASS25.75sNonecsv-money-close
ledger-closerepeat 2 · e014Jev advicePASS23.54scsv-money-close, csv-data-qualitycsv-money-close, csv-data-quality
ledger-closerepeat 2 · e013Astra advicePASS32.13scsv-money-close, csv-data-qualitycsv-money-close, csv-data-quality
ledger-closerepeat 2 · e015No decision callPASS27.41sNonecsv-money-close
support-bundlerepeat 1 · e005Jev advicePASS25.64sconfig-redaction, archive-manifest, json-overlayconfig-redaction, json-overlay
support-bundlerepeat 1 · e004Astra advicePASS27.75sconfig-redaction, json-overlayconfig-redaction, json-overlay
support-bundlerepeat 1 · e006No decision callPASS22.29sNonejson-overlay, config-redaction
support-bundlerepeat 2 · e016Jev advicePASS25.24sconfig-redaction, archive-manifest, json-overlayconfig-redaction, json-overlay
support-bundlerepeat 2 · e018Astra advicePASS26.56sconfig-redaction, json-overlayconfig-redaction, json-overlay
support-bundlerepeat 2 · e017No decision callPASS25.54sNoneconfig-redaction, json-overlay
copy-noterepeat 1 · e007Jev advicePASS12.41sNonearchive-manifest
copy-noterepeat 1 · e009Astra advicePASS15.83sNoneNone
copy-noterepeat 1 · e008No decision callPASS9.46sNoneNone
copy-noterepeat 2 · e021Jev advicePASS13.54sNonearchive-manifest
copy-noterepeat 2 · e020Astra advicePASS11.15sNonearchive-manifest
copy-noterepeat 2 · e019No decision callPASS10.23sNonearchive-manifest
release-notesrepeat 1 · e012Jev advicePASS18.80srelease-note-export, changelog-digest, markdown-tablesrelease-note-export, changelog-digest, markdown-tables
release-notesrepeat 1 · e011Astra advicePASS19.58srelease-note-exportrelease-note-export
release-notesrepeat 1 · e010No decision callPASS18.86sNonerelease-note-export
release-notesrepeat 2 · e023Jev advicePASS18.19srelease-note-export, changelog-digest, markdown-tablesrelease-note-export, changelog-digest, markdown-tables
release-notesrepeat 2 · e022Astra advicePASS22.99srelease-note-exportrelease-note-export
release-notesrepeat 2 · e024No decision callPASS18.22sNonerelease-note-export
Paired timing differences

Positive values mean slower. Each row compares the same task and repetition; run order was shuffled with a fixed seed. Small samples, remote load and provider-cache effects remain.

Task / repeatAstra advice − offJev advice − offJev − Astra advice
t01-ledger-close / 1+0.98s-1.03s-2.01s
t01-ledger-close / 2+4.71s-3.87s-8.59s
t02-support-bundle / 1+5.47s+3.35s-2.12s
t02-support-bundle / 2+1.02s-0.30s-1.32s
t03-copy-note / 1+6.37s+2.95s-3.41s
t03-copy-note / 2+0.92s+3.31s+2.39s
t04-release-notes / 1+0.72s-0.06s-0.78s
t04-release-notes / 2+4.77s-0.02s-4.80s
Shadow checks: advice without agent exposure

Eight separate selector calls on the pre-review snapshot, before either full matrix. Not counted as completed tasks or pooled into main-trial latency.

TaskBackendRecommendedTime
t01-ledger-closeAstra advicecsv-money-close, csv-data-quality3617 ms
t01-ledger-closeJev advicecsv-money-close, csv-data-quality397 ms
t02-support-bundleAstra adviceconfig-redaction, json-overlay3434 ms
t02-support-bundleJev adviceconfig-redaction, json-overlay386 ms
t03-copy-noteAstra adviceNone2349 ms
t03-copy-noteJev adviceNone419 ms
t04-release-notesAstra advicerelease-note-export2565 ms
t04-release-notesJev advicerelease-note-export, changelog-digest, markdown-tables375 ms
Usage, verification and limits

Four task types × three arms × two repetitions = 24 trials. Fresh agent processes and workspaces; identical task prompts, system prompt, tools, ten-skill catalog and non-treatment configuration. Single concurrency, fixed-seed shuffled order (20260919), 12-iteration cap, 150-second run budget and five-second decision deadline. All arms use provider-default reasoning and sampling.

ArmMain uncached inputMain cache readsMain outputAuxiliary usage, native fields
No decision call43,807280,4259,605{}
Astra advice32,837280,4309,992{"prompt_tokens": 21640, "completion_tokens": 1192, "total_tokens": 22832}
Jev advice46,313268,16310,059{"input_tokens": 14872, "output_tokens": 1392}

Token categories are reported as Hermes/provider returned them. Cache reads are not added to uncached input silently. API pricing is unverified: zero cost is not claimed.

  • 1,288 tests passed across 77 selected files after fixing five review findings: gateway image gating, virtual MoA gating, fabricated catalog entries, lossy name decoding and concealed duplicate names. This is not the full repository suite. This matrix ran after the first three fixes; the final two catalog-edge fixes were verified to preserve its exact prompts/candidates/decision payloads. A separate final-source off/LLM/Jev smoke passed 3/3 on the multi-skill support task. The prior matrix is retained, not pooled.
  • Scorer self-check: 11 valid examples accepted, 47 deliberately broken variants rejected. Frozen inputs, exact deliverable sets and output contents checked.
  • Real AIAgent and real tools ran inside bubblewrap. Personal files, scorer, expected outputs and other episodes were not mounted. Provider networking was available; tool-side network prohibition was an instruction, not a network namespace guarantee.
  • Skills and profile configuration were read-only. Credentials arrived through stdin into a scoped in-memory store, not command arguments or environment variables. Raw evidence, credentials, logs and configuration files are not included as downloads.
  • Four synthetic tasks cannot establish general production benefit. Repeats are not additional task diversity. Clear local specifications and a strong main model can create a ceiling effect.
  • The initial full-agent setup failed before tool work because forced reasoning-off was rejected. A separate scorer mount problem was fixed. Both diagnostics and failed attempt are preserved separately, not hidden as quality failures.
  • Publication note: TypeSafe MCA §2.3(f) restricts benchmark publication. The author confirms TypeSafe has permitted publication of these results.

Source: recorded agent-003 episodes, frozen fixtures, shadow-001 and deterministic artifact scores. The Hermes feature remains experimental and off by default; the tested implementation was not installed in the normal runtime. Raw evidence, diagnostics and source patches are not hosted here.

Separate diagnostic, earlier-matrix and smoke results are not pooled into these 24 trials.

Methods and sources

Shared checks, sources and report context

PoCs 1–2: what was verified—and what was not

Real agent execution

Real provider calls, file tools and SQLite persistence. Original task/scorer/source manifests were frozen, artifacts independently re-scored, and each episode's model actions were inspected before continuation.

Not general safety proof

Supervised filesystem isolation, not kernel-enforced network isolation. Correct finite artifacts and reviewed actions do not prove security or broad quality equivalence.

CheckPoC 1PoC 2
Primary artifacts79/8064/64
Inference requests495 main + 32 routing405 main + 16 ranking
Safety receipts8064
Canonical code tests1,256 passed / 78 files600 passed / 22 files
Population8 tasks × 2 repeats × 5 arms8 tasks × 2 repeats × 4 arms
Jev usage: client evidence vs dashboard reconciliation

The saved primary runs contain 16 TypeSafe System One requests each. PoC 2 responses report jev-1.13.0, HTTP 200 and token usage. The adapter targets TypeSafe's documented HTTPS endpoint; the live runner was checked for accidental test transports.

A missing-usage observation in the account dashboard remains unresolved. Historical records lack TypeSafe's response-header request IDs and absolute request timestamps. A separate, later diagnostic through the same adapter and credential returned a provider-issued request ID over real TLS. It confirms the current connection, not independently every historical request or its billing attribution.

The diagnostic is excluded from benchmark counts and timing. Account/key/date-filter reconciliation was not completed. No account identifiers, credential metadata, request IDs or diagnostic files are published here.

Timing, cache and token accounting

Task clocks include auxiliary work, construction and the agent loop. Subprocess startup, scorer time and inter-episode supervision are separate. Request send-to-close spans are SDK/HTTP elapsed, not pure compute. Stream completion and raw iterator exhaustion are distinct; eager calls do not expose TTFT at this boundary.

Provider caching remained enabled; fresh task profiles were not guaranteed cold provider caches. Failed attempts, tool recovery and fallbacks stayed in their denominators. No forced slow baseline, human answer repairs or result-driven reruns were used.

PoC 2 armMain input tokensCached input includedMain output tokens
Ordinary Astra695,203636,13210,861
All-source bundle745,775665,3619,706
Lexical bundle712,021627,51810,536
Jev bundle773,969663,50010,613

Jev ranking added 79,038 input and 13,311 output tokens in PoC 2. The fastest arm used more input tokens than ordinary Astra despite fewer requests. No dollar savings are claimed without verified pricing.

Reproducibility and sources

Original run IDs: routing-primary-001 and retrieval-primary-001. The experimental changes remain uncommitted and uninstalled in Hermes. This page is a report-only publication, not a source release or independent public reproduction.

Model identities: @chatgpt/gpt-6-astra, @chatgpt/gpt-5.6-luna via API; native ranking model jev-1.13.0. Calibration also included @chatgpt/gpt-5.4-mini and @claude/claude-haiku-4-5-20251001.

Results come from retained local execution artifacts and independent audits, not from the documentation links. No raw logs, workspaces, evidence downloads or configuration exports are deployed with this report.

Original report context and study status

Jev returned bounded decisions faster than Astra in PoC 4. That call-level win did not establish an end-to-end task benefit from adding Jev. In the agent pilots, simple source bundling helped; routing, retrieval ranking and skill advice did not demonstrate an incremental Jev benefit.

One report, five separate PoCs. This consolidates the published routing/retrieval findings, backend comparison and full-agent skill-advisor trials. It preserves their distinct populations and timing boundaries; no new inference was run for this consolidation.

PoCs 1, 2, 4 and 5 are completed. Only PoC 3 is on hold. The combined routing-and-retrieval experiment was never run and remains paused until explicitly resumed. Completed pilots retain their findings and limitations; this status update does not start new trials.

PoC 1 / COMPLETEGeneration routingPilot completed · 79/80 artifacts
PoC 2 / COMPLETESource retrievalPilot completed · 64/64 artifacts
PoC 3 / ON HOLDCombined routing + retrievalNever run · no combined result
PoC 4 / COMPLETEJev vs Astra backendPilot completed · 72/72 labels each
PoC 5 / COMPLETEAgent skill advisorPilot completed · 8/8 trials per arm

These are separate populations and mechanisms, not one pooled benchmark. PoC 4 measures decision-call latency; PoCs 1, 2 and 5 measure agent tasks with different timing boundaries. Do not treat their absolute times as a cross-experiment treatment effect.

Recommendation

Keep these Jev policies off by default on this evidence. Judge any future helper against a no-helper baseline, using completed-task quality and time—not just the speed of its own call.

Full recommendation and related reports

Status and recommendation

PoCs 1, 2, 4 and 5 are completed. PoC 3 alone is on hold and has not been run.

The completed pilots remain evidence: simple all-source bundling helped in PoC 2, and Jev cut bounded decision-call latency in PoC 4. Neither finding demonstrates that adding Jev improves end-to-end agent work. PoC 5 found no task-quality gain, and PoC 3 has no result.

Do not enable these Jev routing, retrieval or skill-advisor policies by default on these results. The no-decision-call baseline remains essential to judging whether an extra decision call is useful; a faster advisor is not the same as a faster task.

Only the combined experiment is paused. Completion of the other pilots is not a claim of production readiness, nor a claim that Jev can never help.

Related reading: cross-study Jev/Hermes integration assessment and the separate Hermes Jev Skills pinned benchmark. These are not additional trials in this report.

Report-only publication · September 2026 · PoCs 1, 2, 4 and 5 completed; PoC 3 on hold.
Completed evidence retained · No running Hermes configuration was changed.
← Back to directory