JustHandled Labs
// AI Agents & LLM Ops

Agent Skill Model-Release Regression Harness

Check whether an existing agent skill still improves the same fixtures after a model release without expanding tools, permissions, cost, latency, or output-contract risk.

What problem does Agent Skill Model-Release Regression Harness solve?

A skill that improved output on last month's model can lose its advantage, break an output contract, expand tools or permissions, or become too expensive or slow after a model release while still appearing to work in a single demo.

Use it to

What it returns

A representative input and result

fixture-backed sample
input Compare release-note-evidence-extractor v1.0.0 on example-agent 2026-07 versus 2026-08. For each fixture, supply prior and candidate baseline and skill-assisted runs with score, pass state, output, tools, permissions, cost, latency, and evidence reference.
result Posture: HOLD. extract-dated-facts is CONTRACT_REGRESSION: the candidate assisted run failed, lost the required inferences path, used web_search, requested network permission, dropped skill uplift from +0.37 to -0.09, cost 1.83x, and latency 1.70x. Receipt: 32042E0F...

Access and approval boundaries

Known limitations

Questions

Does this call OpenAI, DeepSeek, Qwen, Anthropic, or another provider?

No. It analyzes recorded observations from a separate compatible runner and requires no API key or model credits.

Why are four observations required per case?

They separate model improvement from skill improvement. A candidate model can score better alone while the skill itself adds less value or makes the result worse.

What happens when the candidate run is missing?

The case remains CANNOT_ASSESS and the report posture is REVIEW unless another supplied case independently requires HOLD.

Does CLEAR prove the skill works on the new model?

Only for the supplied fixtures, outputs, thresholds, and evidence references. It is not a general compatibility or safety certification.

Can it detect tool or permission drift?

Yes. It blocks disallowed tools or permissions and separately flags expansion beyond the prior skill-assisted boundary even when the expanded items were pre-allowed.

Can I use Promptfoo, LangSmith, or my own runner?

Yes. Map the recorded results to the documented JSON contract. The package stays provider-neutral and does not replace those runners.

Does it measure demand for my skill?

No. Run evidence, publication, installs, views, and indexing are operational signals. Qualified buyer demand and revenue require separate external evidence.

Agent Skill Model-Release Regression Harness keeps proof and approval boundaries visible.

The listing includes the tested package, realistic samples, declared permissions, and known limitations.

Get Agent Skill Model-Release Regression Harness on Agensi