EVIDENCE BEFORE PROMISES
A useful workflow has to earn its place.
A polished example is a starting point. Installation checks, behavioral testing, and practitioner review are what turn a prototype into a release.
Current status: prototype collection. No host configuration or paid release has passed the evaluation below. Displayed examples are fictional and hand-authored.
The standard we will test against
Each paid skill needs ten representative cases: six normal, two with missing or conflicting inputs, and two adversarial or out-of-scope. Each case is run three times on each claimed host/model configuration. That is thirty observed runs per configuration.
- Zero observed critical errors, such as invented prices, client facts, or unauthorized commitments.
- At least 90% of required output fields present.
- At least 80% of runs usable with minor edits, judged against the written rubric.
- Installation and first-task checks with five intended users.
These are proposed release thresholds, not current results or performance guarantees. A small sample can reveal failure modes; it cannot prove universal reliability.
Compare against a fair baseline
The packaged workflow and a reasonable general prompt should use the same supplied facts, model, and comparable settings. Review factual fidelity, completeness, correction effort, and task completion. Record the condition, denominator, failure cases, and reviewer. Where practical, the reviewer should not know which output came from which condition.
Make the evidence inspectable
A published compatibility record must identify the package hash, host and surface, model, available version information, operating system, dependencies, test date, reviewer, and results. A vendor supporting the skill format does not establish that our package works well in that host.
Current evidence register
| Item | Current evidence |
|---|---|
| AG-00 package | Authored instructions and templates; manifest and ZIP validation. No behavioral runs. |
| AG-01–04 | Product concepts and fictional reference examples. Paid implementation not released. |
| Host support | Claude Code documentation reviewed; installation and model testing pending. |
| Industry review | Reviewer not yet assigned. No independent certification. |
| Time saved and customer results | Not measured; no claims made. |
When something stops working
An uncertain host variant should be marked for retest and its sales paused. Harmful releases should be withdrawn and affected buyers given a corrected release or a resolution path. That operating workflow must be ready before paid launch.