Huddleston Personal Computer Research

BEARHUDDLESTON.DEV / RESEARCH

Research.

Experiments, what they showed, and what we'd do with the results. Start with the conclusion; open the studies when you want the detail.

AI agents / September 2026

Jev + Hermes

Does an extra decision model make an agent better?

Our conclusion: not enough benefit to enable it by default. Jev answered small questions faster, but adding it did not demonstrate better completed tasks in our tests.

Small, separate studies—not a claim that Jev can never help.

Read the conclusion and recommendation

VerdictKeep it off by default. Faster small decisions; no demonstrated gain in finished work.

Supporting studies
  • Decision Aux pilots

    Our tests of model choice, source selection, decision speed and skill advice. Short findings first; data and methods on demand.

  • Hermes Jev Skills benchmark

    A separate evaluation of the third-party toolkit, including agent tasks, browser workflows and integration defects.

  • Jev approvals: Jev, Mini and Luna

    Replacing an existing reviewer call rather than adding a decision stage: live approval decisions, latency, estimated API cost, and input-loss failures in anpicasso’s plugin.

Four Decision Aux pilots are complete. The combined experiment, PoC 3, remains on hold and was not run.