Best fit
A skill that improved output on last month's model can lose its advantage, break an output contract, expand tools or permissions, or become too expensive or slow after a model release while still appearing to work in a single demo.
Check whether an existing agent skill still improves the same fixtures after a model release without expanding tools, permissions, cost, latency, or output-contract risk.
A skill that improved output on last month's model can lose its advantage, break an output contract, expand tools or permissions, or become too expensive or slow after a model release while still appearing to work in a single demo.
A skill that improved output on last month's model can lose its advantage, break an output contract, expand tools or permissions, or become too expensive or slow after a model release while still appearing to work in a single demo.
CLEAR, REVIEW, or HOLD posture across the supplied release pair. Per-case STABLE, IMPROVED, CANNOT_ASSESS, or regression classification.
It does not call a model, spend credits, or provide provider-specific runner adapters.
No. It analyzes recorded observations from a separate compatible runner and requires no API key or model credits.
They separate model improvement from skill improvement. A candidate model can score better alone while the skill itself adds less value or makes the result worse.
The case remains CANNOT_ASSESS and the report posture is REVIEW unless another supplied case independently requires HOLD.
Only for the supplied fixtures, outputs, thresholds, and evidence references. It is not a general compatibility or safety certification.
Yes. It blocks disallowed tools or permissions and separately flags expansion beyond the prior skill-assisted boundary even when the expanded items were pre-allowed.
Yes. Map the recorded results to the documented JSON contract. The package stays provider-neutral and does not replace those runners.
No. Run evidence, publication, installs, views, and indexing are operational signals. Qualified buyer demand and revenue require separate external evidence.
The listing includes the tested package, realistic samples, declared permissions, and known limitations.
Get Agent Skill Model-Release Regression Harness on Agensi