BEARHUDDLESTON.DEV / DECISION AUX
Decision Aux:
consolidated PoC findings.
Jev made decisions faster, but adding it did not demonstrate better agent work. The useful task-level improvement came from supplying source files without Jev.
Decision Aux is an optional helper step: ask a model which model should do the task, which files it needs, or which skills to load. Jev is one candidate for that helper; Astra does the main work in the source and skill trials.
Just want the recommendation? Read our integration assessment →
Model choice · completed
Did Jev choose the faster task model?
Result. Jev moved 0/16 tasks to Luna, the faster model selected in an earlier screen.
Meaning. It kept Astra every time and added a routing call. This did not demonstrate faster work through model selection.
Data and method: model choice
01 / CHOOSE THE GENERATOR
PoC 1 / COMPLETE · Pilot completed.
Jev kept Astra on every primary task.
The router ran once, before constructing a fresh agent. It could choose the screened faster generator, Luna, or keep Astra. In the Jev arm, every task still ran on Astra.
Seconds per task; lower is faster. Mean shown initially. Each arm has 16 attempts. Times include routing, agent construction and the ordinary tool loop.
| Arm | Artifacts | Mean seconds | Median seconds | Main requests | Aux requests |
|---|---|---|---|---|---|
| Astra only | 16/16 | 19.312 | 19.491 | 103 | 0 |
| Luna only | 15/16 | 16.963 | 14.682 | 98 | 0 |
| Rule router | 16/16 | 18.235 | 17.564 | 99 | 0 |
| Astra router | 16/16 | 19.569 | 20.555 | 96 | 16 |
| Jev router | 16/16 | 18.795 | 19.871 | 99 | 16 |
Each square is one task. Green = Astra; amber = Luna. A router cannot demonstrate faster-generator allocation when it never leaves Astra.
Interpretation: Jev's raw mean was lower than Astra-only, but its median was higher. It added about 0.43 seconds of routing work without allocating any task to Luna. That is not demonstrated generation-routing acceleration.
Quality miss, fallback and uncertainty
- Luna-only passed 15/16. One service-order artifact was incorrect; the model's self-check accepted the same incorrect order. Four model tool errors were retained with model-driven recovery.
- The Astra router timed out twice. Jev had two combined abstained/invalid decisions. All four fell back to Astra; the saved reason does not distinguish Jev abstention from invalidity.
- The rule router also kept Astra in all 16 cases. Its lower mean cannot be credited to faster-generator allocation either.
- Jev saw the request, not the task's specification or source. Cautious policy and insufficient task evidence are possible explanations, not established causes.
| Treatment minus control | Mean paired difference | 95% task-cluster interval |
|---|---|---|
| Luna only − Astra only | -2.349s | -6.272 to +2.911s |
| Rule router − Astra only | -1.076s | -1.917 to -0.235s |
| Astra router − Astra only | +0.258s | -2.060 to +2.252s |
| Jev router − Astra only | -0.516s | -2.095 to +0.809s |
Negative differences favor the named treatment. Eight task clusters, two repeats each; 5,000 clustered bootstrap resamples. Intervals describe task variation here, not every provider/cache effect.
Generator calibration — separate from primary evidence
Three calibration tasks × two repeats × four generators. Eligibility required every attempt to pass, the requested model identity, and real tool execution; lowest mean among eligible models selected Luna. No primary task was used to tune the choice.
| Model via API | Passes | Mean seconds | Eligible |
|---|---|---|---|
| @chatgpt/gpt-6-astra | 6/6 | 15.747 | Yes |
| @chatgpt/gpt-5.4-mini | 5/6 | 15.947 | No |
| @chatgpt/gpt-5.6-luna | 6/6 | 10.067 | Yes |
| @claude/claude-haiku-4-5-20251001 | 6/6 | 19.570 | Yes |
Every recorded primary attempt
| Episode | Task | Repeat | Arm | Artifact | Seconds | Main | Aux |
|---|---|---|---|---|---|---|---|
| e046 | p01 | 1 | Rule router | PASS | 12.034 | 4 | 0 |
| e076 | p01 | 1 | Luna only | PASS | 17.347 | 5 | 0 |
| e042 | p01 | 1 | Jev router | PASS | 12.931 | 4 | 1 |
| e041 | p01 | 1 | Astra router | PASS | 11.382 | 4 | 1 |
| e061 | p01 | 1 | Astra only | PASS | 10.481 | 4 | 0 |
| e040 | p01 | 2 | Rule router | PASS | 10.862 | 4 | 0 |
| e062 | p01 | 2 | Luna only | PASS | 8.890 | 5 | 0 |
| e060 | p01 | 2 | Jev router | PASS | 11.660 | 4 | 1 |
| e033 | p01 | 2 | Astra router | PASS | 10.505 | 4 | 1 |
| e032 | p01 | 2 | Astra only | PASS | 15.721 | 6 | 0 |
| e073 | p02 | 1 | Rule router | PASS | 12.497 | 5 | 0 |
| e057 | p02 | 1 | Luna only | PASS | 10.241 | 5 | 0 |
| e069 | p02 | 1 | Jev router | PASS | 12.202 | 5 | 1 |
| e045 | p02 | 1 | Astra router | PASS | 16.908 | 5 | 1 |
| e025 | p02 | 1 | Astra only | PASS | 11.798 | 5 | 0 |
| e056 | p02 | 2 | Rule router | PASS | 14.872 | 5 | 0 |
| e036 | p02 | 2 | Luna only | PASS | 8.800 | 5 | 0 |
| e053 | p02 | 2 | Jev router | PASS | 14.877 | 5 | 1 |
| e031 | p02 | 2 | Astra router | PASS | 18.330 | 5 | 1 |
| e037 | p02 | 2 | Astra only | PASS | 14.028 | 5 | 0 |
| e019 | p03 | 1 | Rule router | PASS | 14.704 | 5 | 0 |
| e077 | p03 | 1 | Luna only | PASS | 13.085 | 6 | 0 |
| e038 | p03 | 1 | Jev router | PASS | 15.800 | 5 | 1 |
| e043 | p03 | 1 | Astra router | PASS | 14.362 | 5 | 1 |
| e075 | p03 | 1 | Astra only | PASS | 22.694 | 7 | 0 |
| e026 | p03 | 2 | Rule router | PASS | 21.447 | 7 | 0 |
| e058 | p03 | 2 | Luna only | PASS | 16.045 | 7 | 0 |
| e021 | p03 | 2 | Jev router | PASS | 20.878 | 7 | 1 |
| e071 | p03 | 2 | Astra router | PASS | 15.068 | 5 | 1 |
| e008 | p03 | 2 | Astra only | PASS | 19.536 | 7 | 0 |
| e013 | p04 | 1 | Rule router | PASS | 15.960 | 6 | 0 |
| e022 | p04 | 1 | Luna only | FAIL | 14.984 | 5 | 0 |
| e047 | p04 | 1 | Jev router | PASS | 18.123 | 6 | 1 |
| e064 | p04 | 1 | Astra router | PASS | 23.222 | 8 | 1 |
| e015 | p04 | 1 | Astra only | PASS | 16.231 | 6 | 0 |
| e063 | p04 | 2 | Rule router | PASS | 17.855 | 6 | 0 |
| e055 | p04 | 2 | Luna only | PASS | 15.207 | 7 | 0 |
| e059 | p04 | 2 | Jev router | PASS | 16.855 | 6 | 1 |
| e068 | p04 | 2 | Astra router | PASS | 17.992 | 6 | 1 |
| e079 | p04 | 2 | Astra only | PASS | 19.009 | 6 | 0 |
| e035 | p05 | 1 | Rule router | PASS | 23.858 | 7 | 0 |
| e017 | p05 | 1 | Luna only | PASS | 14.380 | 7 | 0 |
| e012 | p05 | 1 | Jev router | PASS | 27.986 | 7 | 1 |
| e067 | p05 | 1 | Astra router | PASS | 22.297 | 6 | 1 |
| e072 | p05 | 1 | Astra only | PASS | 22.267 | 7 | 0 |
| e003 | p05 | 2 | Rule router | PASS | 19.970 | 6 | 0 |
| e009 | p05 | 2 | Luna only | PASS | 24.501 | 8 | 0 |
| e078 | p05 | 2 | Jev router | PASS | 20.295 | 6 | 1 |
| e066 | p05 | 2 | Astra router | PASS | 23.512 | 6 | 1 |
| e052 | p05 | 2 | Astra only | PASS | 24.276 | 6 | 0 |
| e051 | p06 | 1 | Rule router | PASS | 17.273 | 7 | 0 |
| e001 | p06 | 1 | Luna only | PASS | 15.844 | 6 | 0 |
| e010 | p06 | 1 | Jev router | PASS | 20.021 | 7 | 1 |
| e018 | p06 | 1 | Astra router | PASS | 20.109 | 7 | 1 |
| e024 | p06 | 1 | Astra only | PASS | 19.865 | 7 | 0 |
| e011 | p06 | 2 | Rule router | PASS | 17.125 | 7 | 0 |
| e006 | p06 | 2 | Luna only | PASS | 10.079 | 5 | 0 |
| e005 | p06 | 2 | Jev router | PASS | 20.666 | 7 | 1 |
| e014 | p06 | 2 | Astra router | PASS | 21.002 | 7 | 1 |
| e002 | p06 | 2 | Astra only | PASS | 19.447 | 6 | 0 |
| e039 | p07 | 1 | Rule router | PASS | 30.362 | 7 | 0 |
| e080 | p07 | 1 | Luna only | PASS | 41.102 | 8 | 0 |
| e027 | p07 | 1 | Jev router | PASS | 21.856 | 6 | 1 |
| e029 | p07 | 1 | Astra router | PASS | 30.790 | 6 | 1 |
| e020 | p07 | 1 | Astra only | PASS | 26.621 | 7 | 0 |
| e023 | p07 | 2 | Rule router | PASS | 19.907 | 6 | 0 |
| e049 | p07 | 2 | Luna only | PASS | 38.349 | 7 | 0 |
| e054 | p07 | 2 | Jev router | PASS | 19.720 | 6 | 1 |
| e007 | p07 | 2 | Astra router | PASS | 23.500 | 6 | 1 |
| e048 | p07 | 2 | Astra only | PASS | 24.690 | 8 | 0 |
| e044 | p08 | 1 | Rule router | PASS | 19.592 | 8 | 0 |
| e004 | p08 | 1 | Luna only | PASS | 10.866 | 6 | 0 |
| e050 | p08 | 1 | Jev router | PASS | 23.808 | 9 | 1 |
| e034 | p08 | 1 | Astra router | PASS | 22.223 | 8 | 1 |
| e070 | p08 | 1 | Astra only | PASS | 19.215 | 8 | 0 |
| e016 | p08 | 2 | Rule router | PASS | 23.442 | 9 | 0 |
| e028 | p08 | 2 | Luna only | PASS | 11.688 | 6 | 0 |
| e065 | p08 | 2 | Jev router | PASS | 23.049 | 9 | 1 |
| e074 | p08 | 2 | Astra router | PASS | 21.907 | 8 | 1 |
| e030 | p08 | 2 | Astra only | PASS | 23.105 | 8 | 0 |
PoCs 1–2 / the routing and retrieval mechanisms
01 / ROUTETask → choose Astra or Luna → normal agent loop.
02 / RETRIEVESource files → bundle evidence → normal Astra loop.
Source selection · completed
Did Jev help find the right source files?
Result. Giving Astra all the source files, without Jev, cut mean task time by 14.4%. All files fit within the budget.
Meaning. Simple bundling helped here. Jev’s file ranking did not demonstrate an extra benefit over the simpler selection methods.
Data and method: source selection
02 / SUPPLY THE EVIDENCE
PoC 2 / COMPLETE · Pilot completed.
All-source bundling was the useful control.
Astra stayed fixed. Every task had 20 short source files, all within the 9,000-character source budget. One arm supplied all files; two shortlisted up to six using the same import expansion. Ordinary Astra kept unrestricted search, reads and batching.
Seconds per task; lower is faster. Mean shown initially. Ranking, source reads and agent initialization are included.
| Arm | Artifacts | Mean seconds | Median seconds | Main requests | Aux requests |
|---|---|---|---|---|---|
| Ordinary Astra | 16/16 | 17.262 | 17.851 | 110 | 0 |
| All-source bundle | 16/16 | 14.768 | 15.355 | 88 | 0 |
| Lexical bundle | 16/16 | 16.560 | 17.047 | 100 | 0 |
| Jev bundle | 16/16 | 17.103 | 16.669 | 107 | 16 |
14.4% lower mean time
All-source bundling: 14.77 seconds versus ordinary Astra's 17.26 seconds. Every artifact passed. Main requests fell from 110 to 88: a 20% reduction.
No incremental Jev win
Jev averaged 17.10 seconds versus lexical selection's 16.56 seconds. It used 107 main requests plus 16 ranking requests. Its mean retrieval overhead was about 0.51 seconds per task.
| Treatment minus control | Mean paired difference | 95% task-cluster interval |
|---|---|---|
| All-source bundle − Ordinary Astra | -2.494s | -4.285 to -0.743s |
| Lexical bundle − Ordinary Astra | -0.702s | -2.381 to +1.040s |
| Jev bundle − Ordinary Astra | -0.159s | -1.768 to +1.249s |
| Jev bundle − Lexical bundle | +0.543s | -1.100 to +2.239s |
| Jev bundle − All-source bundle | +2.335s | +1.306 to +3.640s |
Task-cluster bootstrap intervals; negative means faster. Jev–lexical spans both benefit and harm. Every task-family mean favored all-source bundling over Jev in this run. This does not establish general quality equivalence.
Fewer reads did not guarantee fewer rounds.
| Arm | Read calls | Search calls | Reads of bundled files | Mean injected characters |
|---|---|---|---|---|
| Ordinary Astra | 152 | 16 | 0 | 0 |
| All-source bundle | 87 | 1 | 55 | 8040 |
| Lexical bundle | 136 | 12 | 0 | 3190 |
| Jev bundle | 119 | 8 | 62 | 3067 |
Jev supplied governing code that the lexical shortlist often missed. Astra nevertheless reread bundled files 62 times. Those reads can be legitimate verification; fewer search/read calls alone did not produce fewer inference requests.
All-source: 3 main requests, 8.81s. Jev: 7 main requests, 18.30s. This single example illustrates the mechanism, not a separate aggregate claim.
Method limits and task-family differences
- The source trees were small: all 20 files fit. This is not a large-repository RAG benchmark.
- Lexical selection used BM25 over the full task request. Shared instruction wording dominated many selections, favoring README and operational prose. It is not evidence against a tuned, task-focused retrieval baseline. No after-the-fact retuning was performed.
- All-source bundles averaged 4,800 source characters. Framing increased the injected text to about 8,040 characters. The six-file Jev bundle averaged about 3,067 injected characters.
- Jev ranked successfully in all 16 primary episodes. No provider, tool, interruption or fallback errors, repairs or reruns were recorded.
- Exact source text was appended once to the initial user request. The system prefix, toolset and Astra identity were unchanged; original sources remained accessible.
| Task family | All-source − ordinary | Lexical − ordinary | Jev − ordinary |
|---|---|---|---|
| r01 | -0.63s | +0.45s | +1.25s |
| r02 | +0.70s | +0.13s | +1.84s |
| r03 | -5.45s | -1.88s | -4.12s |
| r04 | -3.85s | -2.92s | -0.36s |
| r05 | -3.63s | +0.08s | -3.02s |
| r06 | -6.42s | -5.15s | -0.39s |
| r07 | -1.00s | +3.71s | +2.33s |
| r08 | +0.33s | -0.03s | +1.20s |
Every recorded primary attempt
| Episode | Task | Repeat | Arm | Artifact | Seconds | Main | Aux |
|---|---|---|---|---|---|---|---|
| e050 | r01 | 1 | All-source bundle | PASS | 12.787 | 5 | 0 |
| e051 | r01 | 1 | Lexical bundle | PASS | 15.641 | 6 | 0 |
| e028 | r01 | 1 | Jev bundle | PASS | 14.194 | 6 | 1 |
| e056 | r01 | 1 | Ordinary Astra | PASS | 15.062 | 7 | 0 |
| e063 | r01 | 2 | All-source bundle | PASS | 12.181 | 5 | 0 |
| e036 | r01 | 2 | Lexical bundle | PASS | 11.478 | 5 | 0 |
| e006 | r01 | 2 | Jev bundle | PASS | 14.520 | 6 | 1 |
| e019 | r01 | 2 | Ordinary Astra | PASS | 11.157 | 4 | 0 |
| e009 | r02 | 1 | All-source bundle | PASS | 16.123 | 6 | 0 |
| e027 | r02 | 1 | Lexical bundle | PASS | 15.248 | 7 | 0 |
| e064 | r02 | 1 | Jev bundle | PASS | 14.934 | 6 | 1 |
| e011 | r02 | 1 | Ordinary Astra | PASS | 15.966 | 7 | 0 |
| e054 | r02 | 2 | All-source bundle | PASS | 14.939 | 6 | 0 |
| e025 | r02 | 2 | Lexical bundle | PASS | 14.672 | 6 | 0 |
| e029 | r02 | 2 | Jev bundle | PASS | 18.413 | 7 | 1 |
| e003 | r02 | 2 | Ordinary Astra | PASS | 13.694 | 5 | 0 |
| e026 | r03 | 1 | All-source bundle | PASS | 16.097 | 6 | 0 |
| e021 | r03 | 1 | Lexical bundle | PASS | 20.103 | 8 | 0 |
| e001 | r03 | 1 | Jev bundle | PASS | 15.900 | 6 | 1 |
| e032 | r03 | 1 | Ordinary Astra | PASS | 20.611 | 8 | 0 |
| e035 | r03 | 2 | All-source bundle | PASS | 15.380 | 5 | 0 |
| e005 | r03 | 2 | Lexical bundle | PASS | 18.510 | 7 | 0 |
| e014 | r03 | 2 | Jev bundle | PASS | 18.245 | 7 | 1 |
| e010 | r03 | 2 | Ordinary Astra | PASS | 21.767 | 8 | 0 |
| e022 | r04 | 1 | All-source bundle | PASS | 14.773 | 5 | 0 |
| e047 | r04 | 1 | Lexical bundle | PASS | 21.539 | 8 | 0 |
| e020 | r04 | 1 | Jev bundle | PASS | 21.396 | 7 | 1 |
| e058 | r04 | 1 | Ordinary Astra | PASS | 20.170 | 8 | 0 |
| e053 | r04 | 2 | All-source bundle | PASS | 15.331 | 6 | 0 |
| e030 | r04 | 2 | Lexical bundle | PASS | 10.423 | 4 | 0 |
| e049 | r04 | 2 | Jev bundle | PASS | 15.688 | 7 | 1 |
| e037 | r04 | 2 | Ordinary Astra | PASS | 17.641 | 8 | 0 |
| e052 | r05 | 1 | All-source bundle | PASS | 15.276 | 6 | 0 |
| e015 | r05 | 1 | Lexical bundle | PASS | 21.094 | 8 | 0 |
| e031 | r05 | 1 | Jev bundle | PASS | 16.849 | 7 | 1 |
| e008 | r05 | 1 | Ordinary Astra | PASS | 18.988 | 8 | 0 |
| e007 | r05 | 2 | All-source bundle | PASS | 16.846 | 6 | 0 |
| e046 | r05 | 2 | Lexical bundle | PASS | 18.453 | 7 | 0 |
| e062 | r05 | 2 | Jev bundle | PASS | 16.488 | 6 | 1 |
| e048 | r05 | 2 | Ordinary Astra | PASS | 20.392 | 8 | 0 |
| e024 | r06 | 1 | All-source bundle | PASS | 8.811 | 3 | 0 |
| e033 | r06 | 1 | Lexical bundle | PASS | 12.813 | 4 | 0 |
| e055 | r06 | 1 | Jev bundle | PASS | 18.305 | 7 | 1 |
| e061 | r06 | 1 | Ordinary Astra | PASS | 16.830 | 6 | 0 |
| e012 | r06 | 2 | All-source bundle | PASS | 13.240 | 5 | 0 |
| e038 | r06 | 2 | Lexical bundle | PASS | 11.767 | 4 | 0 |
| e016 | r06 | 2 | Jev bundle | PASS | 15.810 | 7 | 1 |
| e034 | r06 | 2 | Ordinary Astra | PASS | 18.061 | 7 | 0 |
| e057 | r07 | 1 | All-source bundle | PASS | 15.392 | 6 | 0 |
| e018 | r07 | 1 | Lexical bundle | PASS | 22.488 | 8 | 0 |
| e042 | r07 | 1 | Jev bundle | PASS | 17.161 | 7 | 1 |
| e040 | r07 | 1 | Ordinary Astra | PASS | 20.384 | 8 | 0 |
| e041 | r07 | 2 | All-source bundle | PASS | 16.718 | 6 | 0 |
| e045 | r07 | 2 | Lexical bundle | PASS | 19.054 | 7 | 0 |
| e059 | r07 | 2 | Jev bundle | PASS | 21.619 | 8 | 1 |
| e004 | r07 | 2 | Ordinary Astra | PASS | 13.733 | 5 | 0 |
| e017 | r08 | 1 | All-source bundle | PASS | 16.764 | 6 | 0 |
| e002 | r08 | 1 | Lexical bundle | PASS | 13.007 | 4 | 0 |
| e013 | r08 | 1 | Jev bundle | PASS | 18.269 | 7 | 1 |
| e060 | r08 | 1 | Ordinary Astra | PASS | 13.287 | 6 | 0 |
| e039 | r08 | 2 | All-source bundle | PASS | 15.632 | 6 | 0 |
| e023 | r08 | 2 | Lexical bundle | PASS | 18.671 | 7 | 0 |
| e044 | r08 | 2 | Jev bundle | PASS | 15.862 | 6 | 1 |
| e043 | r08 | 2 | Ordinary Astra | PASS | 18.450 | 7 | 0 |
Combined test · on hold
Did model choice and source selection help together?
Result. This experiment was never run.
Meaning. The separate studies cannot establish a combined speedup. There is no result to interpret.
Why there is no combined result
03 / COMBINE ROUTING AND RETRIEVAL
PoC 3 / ON HOLD · Never run.
No combined result.
The neither/routing/retrieval/both comparison has not been run on the reserved common holdout. PoCs 1 and 2 cannot be combined after the fact to claim synergy or a combined speedup. This experiment remains paused.
Decision speed · completed
Could Jev answer the helper call faster?
Result. Jev returned the same correct labels as Astra with 72.3% less median call time in this test.
Meaning. That is a real call-level latency win, not evidence of better answers or faster completed tasks. The runs used different APIs at different times.
Data and method: decision speed
04 / JEV VS ASTRA BACKEND
PoC 4 / COMPLETE · Pilot completed.
Jev returned the same labels sooner.
The backend comparison used 72 unique, clear synthetic cases. Both Jev and Astra returned all 72/72 labels correctly. Each backend made 102 calls including CLI, warmup, repeat and padding controls; those controls are not additional independent accuracy samples.
Separate sequential runs at different times, different APIs and client behavior, and provider-default sampling for Astra limit the comparison. Timings include network and validation, not just model compute. No verified dollar comparison is available.
Median decision-call latency
Shared scale: 0–3,000 ms. The source report records 72.3% less call time for Jev, not a task-completion speedup. Primary latency uses the 72 unique cases.
Call-level win, not a quality win. The corpus saturated both backends. It establishes neither a quality advantage nor an end-to-end agent benefit. It does show that Jev can make these bounded decision calls faster.
Jev 1.13.0 versus your configured Astra model via API. This compares the two Decision Aux backend paths—not autonomous-agent performance with the feature switched on and off.
What this supports: Jev reduced latency on these bounded judgments. What it does not support: a quality advantage, production accuracy claim, or end-task improvement.
Measured comparison
| Metric | Jev | No Jev: Astra |
|---|---|---|
| Valid responses, final run | 102 / 102 | 102 / 102 |
| Unique labels correct | 72 / 72 | 72 / 72 |
| Median runtime call | 375 ms | 1350 ms |
| P95 runtime call | 464 ms | 1606 ms |
| Maximum runtime call | 534 ms | 2667 ms |
| Fresh CLI process range · n=3 | 732–856 ms | 2184–3378 ms |
| Repeat label agreement | 12 / 12 | 12 / 12 |
| Short-padding label agreement | 12 / 12 | 12 / 12 |
| Input / prompt tokens · all 102 | 56,429 | 52,201 |
| Output / completion tokens · all 102 | 2,788 | 2,424 |
| Estimated cost | $0.00237 | Not verified |
Primary latency uses only the 72 unique cases. The other 30 calls are CLI/warmup/repeat/padding controls, not additional independent accuracy samples. Different tokenizers and response protocols: token counts are not a cost ratio.
The latency distributions
Astra's median was 3.60× Jev's. Jev was faster on all 72 paired inputs; median paired difference 955 ms. These were separate sequential runs, so time, service load and routing can confound the comparison.
Quality: a tie on every task
Both had zero false acceptances among 12 unsupported-evidence examples, and zero severity errors among 24 examples. The corpus is too easy to distinguish judgment quality. No probabilities or confidence were fabricated for the LLM backend.
The comparison found a real compatibility bug
The initial Astra smoke failed: the relay returned HTTP 400, upstream_unavailable, “Model unavailable.” A plain request worked. Single-variable probes showed that temperature: 0 caused the rejection; strict JSON schema without explicit temperature worked.
Fixed the new LLM backend to request the provider's default sampling setting (temperature=None). Kept strict schema and local validation. Expected labels, task instructions, Jev code, timeout and schedule were unchanged. Two regression checks failed before the fix; afterwards 874 tests passed across 58 selected files.
The failed smoke remains in llm-v2-001. Nine diagnostic inference attempts, including five successes, are excluded from benchmark scores and timings; their known usage and failures are included in the evidence. This is not a claim that the first unmodified baseline run succeeded.
Fairness, timing and limits
- Backend E2E comparison, not full autonomous-agent on/off A/B.
- Clear synthetic corpus saturates both models: no quality advantage established.
- Same corpus and schedule; separate runs at different times, not interleaved.
- Jev request-scoped HTTP clients versus host-cached LLM client; network/routes differ.
- Labels and task rubrics identical; each adapter uses its native request/response protocol.
- Astra uses provider-default sampling after compatibility fix; schema and local validation remain enabled.
- Only 72 distinct cases per backend; repeated phases do not increase sample independence.
- Five-second timeout and concurrency one; no claim about load behavior or production tails.
- No verified dollar comparison without API pricing.
- TypeSafe MCA restricts benchmark publication.
The ordinary LLM backend uses Hermes's cached client; the Jev adapter builds a request-scoped client. Each native adapter translates the same bounded questions into its own protocol. Timings include network and validation, not just GPU/model time. Initial plugin discovery is excluded from runtime timings but included in fresh CLI wall time.
Observed all 99 in-process baseline calls: one HTTP response per call, no recovery-log events. The three CLI calls exited successfully but their internal request count was not instrumented. Backend credentials were isolated; the no-Jev profile contained no Jev key or plugin.
Model: @chatgpt/gpt-6-astra via API. This is the configured main-model baseline, not a small optimized classification model. No API price was established, so no dollar-savings ratio is reported.
All phase timing populations
| Phase | N | Jev median ms | Astra median ms | Jev range ms | Astra range ms |
|---|---|---|---|---|---|
| Fresh CLI processes | 3 each | 751.34 | 2354.29 | 732–856 | 2184–3378 |
| Initial in-process calls | 3 each | 408.94 | 1448.10 | 372–443 | 1275–2100 |
| Unique cases | 72 each | 374.65 | 1350.49 | 324–534 | 1160–2667 |
| Repeat controls | 12 each | 368.64 | 1322.89 | 354–498 | 1210–2011 |
| Short padding controls | 12 each | 366.54 | 1337.14 | 331–523 | 1197–1616 |
Small CLI/warmup samples do not establish reliable tail latency. Padding is roughly 1.5k tokens, not a broad long-context test.
Inspect all 72 paired labels and latencies
| Case | Expected | Jev | Astra | Jev ms | Astra ms |
|---|---|---|---|---|---|
| evidence-01 | True | True | True | 389.5 | 1269.0 |
| evidence-02 | True | True | True | 345.3 | 1427.7 |
| evidence-03 | True | True | True | 323.5 | 1340.3 |
| evidence-04 | True | True | True | 387.4 | 1396.0 |
| evidence-05 | True | True | True | 338.4 | 1366.5 |
| evidence-06 | True | True | True | 437.2 | 1315.3 |
| evidence-07 | True | True | True | 344.2 | 1291.5 |
| evidence-08 | True | True | True | 394.6 | 1294.8 |
| evidence-09 | True | True | True | 367.5 | 1370.8 |
| evidence-10 | True | True | True | 374.3 | 1284.9 |
| evidence-11 | True | True | True | 344.9 | 1459.3 |
| evidence-12 | True | True | True | 334.1 | 1402.3 |
| evidence-13 | False | False | False | 398.8 | 1301.0 |
| evidence-14 | False | False | False | 399.6 | 1294.6 |
| evidence-15 | False | False | False | 383.6 | 1442.5 |
| evidence-16 | False | False | False | 374.7 | 1219.0 |
| evidence-17 | False | False | False | 324.4 | 1445.5 |
| evidence-18 | False | False | False | 356.3 | 1570.6 |
| evidence-19 | False | False | False | 362.2 | 1412.3 |
| evidence-20 | False | False | False | 390.0 | 1439.1 |
| evidence-21 | False | False | False | 395.0 | 1316.1 |
| evidence-22 | False | False | False | 433.5 | 1307.1 |
| evidence-23 | False | False | False | 388.0 | 1240.0 |
| evidence-24 | False | False | False | 385.3 | 2667.4 |
| severity-01 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 450.3 | 1510.4 |
| severity-02 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 406.9 | 1346.3 |
| severity-03 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 418.8 | 1352.4 |
| severity-04 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 343.2 | 1397.9 |
| severity-05 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 406.0 | 1352.8 |
| severity-06 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 366.1 | 1288.8 |
| severity-07 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 352.2 | 1227.2 |
| severity-08 | Cosmetic only; no functional impact | Cosmetic only; no functional impact | Cosmetic only; no functional impact | 334.1 | 1247.8 |
| severity-09 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 483.9 | 1404.6 |
| severity-10 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 396.3 | 1391.7 |
| severity-11 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 436.1 | 1982.7 |
| severity-12 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 358.3 | 1317.4 |
| severity-13 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 484.0 | 1247.1 |
| severity-14 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 449.4 | 1470.4 |
| severity-15 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 383.6 | 1506.3 |
| severity-16 | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | Function degraded or unavailable, but a practical workaround exists | 355.0 | 1278.1 |
| severity-17 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 339.6 | 1395.4 |
| severity-18 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 376.2 | 1442.8 |
| severity-19 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 356.1 | 1405.6 |
| severity-20 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 357.6 | 1299.8 |
| severity-21 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 370.9 | 1593.5 |
| severity-22 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 460.9 | 1423.9 |
| severity-23 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 371.6 | 1460.0 |
| severity-24 | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | Core function blocked without a workaround, or permanent data loss | 388.7 | 1340.6 |
| skill-01 | github | github | github | 343.8 | 1222.2 |
| skill-02 | github | github | github | 349.9 | 1160.0 |
| skill-03 | github | github | github | 391.3 | 1423.9 |
| skill-04 | github | github | github | 534.1 | 1620.3 |
| skill-05 | github | github | github | 409.4 | 1445.2 |
| skill-06 | github | github | github | 467.3 | 1267.6 |
| skill-07 | 371.9 | 1271.2 | |||
| skill-08 | 354.2 | 1269.3 | |||
| skill-09 | 336.8 | 1425.9 | |||
| skill-10 | 375.6 | 1401.0 | |||
| skill-11 | 358.9 | 1282.2 | |||
| skill-12 | 397.0 | 1345.4 | |||
| skill-13 | xlsx | xlsx | xlsx | 378.1 | 1414.4 |
| skill-14 | xlsx | xlsx | xlsx | 419.4 | 1210.8 |
| skill-15 | xlsx | xlsx | xlsx | 374.6 | 1164.8 |
| skill-16 | xlsx | xlsx | xlsx | 367.4 | 1348.6 |
| skill-17 | xlsx | xlsx | xlsx | 366.3 | 1308.6 |
| skill-18 | xlsx | xlsx | xlsx | 330.2 | 1329.6 |
| skill-19 | none | none | none | 374.4 | 1295.2 |
| skill-20 | none | none | none | 354.6 | 1291.6 |
| skill-21 | none | none | none | 371.7 | 1289.0 |
| skill-22 | none | none | none | 416.6 | 1356.4 |
| skill-23 | none | none | none | 387.4 | 1984.8 |
| skill-24 | none | none | none | 341.8 | 1445.1 |
Reproducibility and privacy
Corpus SHA-256: 2bdea01789121cf951d2b821202b342aca18d78154597f62d78b063105d7168f.
Jev run: live-v2-001.
No-Jev final run: llm-v2-002.
This report-only edition includes the comparison charts, methodology and paired-case table. Raw result files, manifests, the separate CSV, comparison script, diagnostic files, test output and source patch are not included as downloads.
Jev estimated cost uses reported input usage and TypeSafe's published $0.042/M input-token price, not invoice billing. TypeSafe MCA §2.3 restricts benchmark publication.
Measured backend E2E comparison. No fabricated model output. Normal Hermes installation and profile settings unchanged. Report checks cover DOM/layout, not pixel inspection.
Skill advice · completed
Did faster skill advice improve the finished work?
Result. Every group completed 8/8 trials, including the group with no helper call.
Meaning. Jev advised faster than Astra, but no task-quality gain was demonstrated. More skill loading was not an improvement by itself.
Data and method: skill advice
05 / AGENT SKILL ADVISOR
PoC 5 / COMPLETE · Pilot completed.
A faster advisor did not establish a better task outcome.
The full-agent pilot ran 24 trials: four tasks × two repeats × three arms, with Astra as the main model throughout. No decision call, Astra advice and Jev advice each had 8/8 artifact-complete trials. No task-quality gain was demonstrated.
Four synthetic tasks and a ten-skill catalog do not establish production benefit or large-catalog recall. Repeats add no task diversity. Recommendation and actual skill loading are separate. The no-decision-call baseline is essential: comparing only two advisors would miss whether advice helped at all.
Median full-agent task duration
Shared scale: 0–40s. Includes the decision call, main model, tools and finalization; excludes initialization. These mixed-task medians are descriptive, not a causal speed estimate.
Advisor medians were 2870 ms for Astra and 395 ms for Jev. Total skill loads were 9 with no decision call, 11 with Astra advice and 16 with Jev advice. More loading was not a quality improvement.
Main model: @chatgpt/gpt-6-astra via API in every arm. Jev backend: jev-1.13.0.
One main model. The same four tasks. Three decision paths. Each trial ends in real files checked by deterministic code—not a model grading itself.
All arms completed every artifact correctly; this pilot did not demonstrate a task-quality gain from the extra call. The no-decision baseline is part of every future E2E comparison.
Jev advice: 16 loads · Astra advice: 11 · No decision call: 9. In the release task, Jev recommended an internal-changelog recipe and a Markdown-table recipe; the agent loaded those alongside the release exporter. The requested artifacts still passed, but advice added unnecessary workflow selection. This ten-skill catalog does not test large-catalog recall.
Outcomes, advisor work and tool usage
No decision call
artifact-complete trials
- Task median
- 20.57s
- Advisor median
- —
- Skill loads
- 9
- Tool calls
- 61
Astra advice
artifact-complete trials
- Task median
- 24.77s
- Advisor median
- 2870 ms
- Skill loads
- 11
- Tool calls
- 63
Jev advice
artifact-complete trials
- Task median
- 21.17s
- Advisor median
- 395 ms
- Skill loads
- 16
- Tool calls
- 68
Complete task duration
Every dot is one trial. Eight per arm; four task types repeated twice. White ticks mark medians.
Includes decision call + main model + tools + finalization. Excludes initialization. Mixed task durations are descriptive, not a causal speed estimate.
What happened in each task?
Recommendation and actual skill loading are separate. An agent can ignore advice or load additional skills.
24 trials shown.
| Task / repeat | Arm | Artifacts | Time | Recommended | Actually loaded |
|---|---|---|---|---|---|
| ledger-closerepeat 1 · e003 | Jev advice | PASS | 24.72s | csv-money-close, csv-data-quality | csv-money-close, csv-data-quality |
| ledger-closerepeat 1 · e002 | Astra advice | PASS | 26.72s | csv-money-close, csv-data-quality | csv-money-close, csv-data-quality |
| ledger-closerepeat 1 · e001 | No decision call | PASS | 25.75s | None | csv-money-close |
| ledger-closerepeat 2 · e014 | Jev advice | PASS | 23.54s | csv-money-close, csv-data-quality | csv-money-close, csv-data-quality |
| ledger-closerepeat 2 · e013 | Astra advice | PASS | 32.13s | csv-money-close, csv-data-quality | csv-money-close, csv-data-quality |
| ledger-closerepeat 2 · e015 | No decision call | PASS | 27.41s | None | csv-money-close |
| support-bundlerepeat 1 · e005 | Jev advice | PASS | 25.64s | config-redaction, archive-manifest, json-overlay | config-redaction, json-overlay |
| support-bundlerepeat 1 · e004 | Astra advice | PASS | 27.75s | config-redaction, json-overlay | config-redaction, json-overlay |
| support-bundlerepeat 1 · e006 | No decision call | PASS | 22.29s | None | json-overlay, config-redaction |
| support-bundlerepeat 2 · e016 | Jev advice | PASS | 25.24s | config-redaction, archive-manifest, json-overlay | config-redaction, json-overlay |
| support-bundlerepeat 2 · e018 | Astra advice | PASS | 26.56s | config-redaction, json-overlay | config-redaction, json-overlay |
| support-bundlerepeat 2 · e017 | No decision call | PASS | 25.54s | None | config-redaction, json-overlay |
| copy-noterepeat 1 · e007 | Jev advice | PASS | 12.41s | None | archive-manifest |
| copy-noterepeat 1 · e009 | Astra advice | PASS | 15.83s | None | None |
| copy-noterepeat 1 · e008 | No decision call | PASS | 9.46s | None | None |
| copy-noterepeat 2 · e021 | Jev advice | PASS | 13.54s | None | archive-manifest |
| copy-noterepeat 2 · e020 | Astra advice | PASS | 11.15s | None | archive-manifest |
| copy-noterepeat 2 · e019 | No decision call | PASS | 10.23s | None | archive-manifest |
| release-notesrepeat 1 · e012 | Jev advice | PASS | 18.80s | release-note-export, changelog-digest, markdown-tables | release-note-export, changelog-digest, markdown-tables |
| release-notesrepeat 1 · e011 | Astra advice | PASS | 19.58s | release-note-export | release-note-export |
| release-notesrepeat 1 · e010 | No decision call | PASS | 18.86s | None | release-note-export |
| release-notesrepeat 2 · e023 | Jev advice | PASS | 18.19s | release-note-export, changelog-digest, markdown-tables | release-note-export, changelog-digest, markdown-tables |
| release-notesrepeat 2 · e022 | Astra advice | PASS | 22.99s | release-note-export | release-note-export |
| release-notesrepeat 2 · e024 | No decision call | PASS | 18.22s | None | release-note-export |
Paired timing differences
Positive values mean slower. Each row compares the same task and repetition; run order was shuffled with a fixed seed. Small samples, remote load and provider-cache effects remain.
| Task / repeat | Astra advice − off | Jev advice − off | Jev − Astra advice |
|---|---|---|---|
| t01-ledger-close / 1 | +0.98s | -1.03s | -2.01s |
| t01-ledger-close / 2 | +4.71s | -3.87s | -8.59s |
| t02-support-bundle / 1 | +5.47s | +3.35s | -2.12s |
| t02-support-bundle / 2 | +1.02s | -0.30s | -1.32s |
| t03-copy-note / 1 | +6.37s | +2.95s | -3.41s |
| t03-copy-note / 2 | +0.92s | +3.31s | +2.39s |
| t04-release-notes / 1 | +0.72s | -0.06s | -0.78s |
| t04-release-notes / 2 | +4.77s | -0.02s | -4.80s |
Shadow checks: advice without agent exposure
Eight separate selector calls on the pre-review snapshot, before either full matrix. Not counted as completed tasks or pooled into main-trial latency.
| Task | Backend | Recommended | Time |
|---|---|---|---|
| t01-ledger-close | Astra advice | csv-money-close, csv-data-quality | 3617 ms |
| t01-ledger-close | Jev advice | csv-money-close, csv-data-quality | 397 ms |
| t02-support-bundle | Astra advice | config-redaction, json-overlay | 3434 ms |
| t02-support-bundle | Jev advice | config-redaction, json-overlay | 386 ms |
| t03-copy-note | Astra advice | None | 2349 ms |
| t03-copy-note | Jev advice | None | 419 ms |
| t04-release-notes | Astra advice | release-note-export | 2565 ms |
| t04-release-notes | Jev advice | release-note-export, changelog-digest, markdown-tables | 375 ms |
Usage, verification and limits
Four task types × three arms × two repetitions = 24 trials. Fresh agent processes and workspaces; identical task prompts, system prompt, tools, ten-skill catalog and non-treatment configuration. Single concurrency, fixed-seed shuffled order (20260919), 12-iteration cap, 150-second run budget and five-second decision deadline. All arms use provider-default reasoning and sampling.
| Arm | Main uncached input | Main cache reads | Main output | Auxiliary usage, native fields |
|---|---|---|---|---|
| No decision call | 43,807 | 280,425 | 9,605 | {} |
| Astra advice | 32,837 | 280,430 | 9,992 | {"prompt_tokens": 21640, "completion_tokens": 1192, "total_tokens": 22832} |
| Jev advice | 46,313 | 268,163 | 10,059 | {"input_tokens": 14872, "output_tokens": 1392} |
Token categories are reported as Hermes/provider returned them. Cache reads are not added to uncached input silently. API pricing is unverified: zero cost is not claimed.
- 1,288 tests passed across 77 selected files after fixing five review findings: gateway image gating, virtual MoA gating, fabricated catalog entries, lossy name decoding and concealed duplicate names. This is not the full repository suite. This matrix ran after the first three fixes; the final two catalog-edge fixes were verified to preserve its exact prompts/candidates/decision payloads. A separate final-source off/LLM/Jev smoke passed 3/3 on the multi-skill support task. The prior matrix is retained, not pooled.
- Scorer self-check: 11 valid examples accepted, 47 deliberately broken variants rejected. Frozen inputs, exact deliverable sets and output contents checked.
- Real AIAgent and real tools ran inside bubblewrap. Personal files, scorer, expected outputs and other episodes were not mounted. Provider networking was available; tool-side network prohibition was an instruction, not a network namespace guarantee.
- Skills and profile configuration were read-only. Credentials arrived through stdin into a scoped in-memory store, not command arguments or environment variables. Raw evidence, credentials, logs and configuration files are not included as downloads.
- Four synthetic tasks cannot establish general production benefit. Repeats are not additional task diversity. Clear local specifications and a strong main model can create a ceiling effect.
- The initial full-agent setup failed before tool work because forced reasoning-off was rejected. A separate scorer mount problem was fixed. Both diagnostics and failed attempt are preserved separately, not hidden as quality failures.
- Publication note: TypeSafe MCA §2.3(f) restricts benchmark publication. The author confirms TypeSafe has permitted publication of these results.
Source: recorded agent-003 episodes, frozen fixtures, shadow-001 and deterministic artifact scores. The Hermes feature remains experimental and off by default; the tested implementation was not installed in the normal runtime. Raw evidence, diagnostics and source patches are not hosted here.
Separate diagnostic, earlier-matrix and smoke results are not pooled into these 24 trials.
Methods and sources
Recommendation
Keep these Jev policies off by default on this evidence. Judge any future helper against a no-helper baseline, using completed-task quality and time—not just the speed of its own call.
Full recommendation and related reports
Status and recommendation
PoCs 1, 2, 4 and 5 are completed. PoC 3 alone is on hold and has not been run.
The completed pilots remain evidence: simple all-source bundling helped in PoC 2, and Jev cut bounded decision-call latency in PoC 4. Neither finding demonstrates that adding Jev improves end-to-end agent work. PoC 5 found no task-quality gain, and PoC 3 has no result.
Do not enable these Jev routing, retrieval or skill-advisor policies by default on these results. The no-decision-call baseline remains essential to judging whether an extra decision call is useful; a faster advisor is not the same as a faster task.
Only the combined experiment is paused. Completion of the other pilots is not a claim of production readiness, nor a claim that Jev can never help.
Related reading: cross-study Jev/Hermes integration assessment and the separate Hermes Jev Skills pinned benchmark. These are not additional trials in this report.