What the runs say, computed not guessed
improve as evidence - a proposal is only as honest as the numbers under it.
Dry runs are counted separately and excluded from every rate: a rehearsal never calls a write tool, so counting one in a tool’s error rate would be counting a call that never happened.
Repairs px0 will make itself
px0 changes revert undoes it.
Telling px0 when the output was wrong
The one thing no record can infer is whether what a run produced was any good. A digest that runs green every week and comes back useless looks perfect in every field px0 has.px0 runs as [bad]. It is what the next step actually learns from - and what px0 memory suggest reads for standing facts.
Having px0 revise the workflow
px0 workflows edit performs.
Three rules hold, each because the obvious alternative is worse:
- Nothing is applied without being shown. An improvement pass that quietly rewrote a scheduled workflow would be the one place px0 stopped listing what it was about to do.
- Tools are never widened on a model’s say-so. A proposal may argue for a new tool, and that argument is printed, but the tool only arrives through the same confirm-and-authorize path
px0 workflows newuses. - A complaint about form belongs in a guideline. “The summary is too long” is not a fact about one workflow; it is a standard. Guideline edits are confirmed separately and appended - your own wording above them is never touched.
--dry-run shows a proposal and applies none of it. --show-evidence prints exactly what the model would be given, as JSON, and makes no model call at all - if you disagree with a proposal, that is what it was reasoning over.
Checking a revision before trusting it
A revision used to mean waiting until Friday to find out whether it helped. The reason it could not be checked sooner is that a workflow run twice compares two different worlds - the pull requests moved, the inbox filled. Let a workflow keep what it read:px0 workflows improve offers this in line, where a fixture exists, before you are asked to accept anything.
Neither the input tools nor the run’s own tools are called during a replay. The comparison is about what a workflow says; letting it act would both change the world and put back the variance the fixture removes.
One fixture is one data point. Replay a second captured run before trusting a difference.
When px0 gives up
An unattended workflow that fails the same wayruns.disable_after_failures times in a row (5 by default) is parked, and you are told through the same channel failures use. A dead connector otherwise means an hourly failure and an hourly notification for the rest of the week, with nothing learning that nothing has changed.
It requires the same cause each time - a workflow failing three different ways is one to look at, not one stuck. A manual run never trips it: you are there, reading the error. The park is a versioned change, and px0 workflows enable is the ordinary way back.

