A portable judgment layer for tool-calling agents

Build the right thing.

Purpose Check helps capable agents notice when execution is drifting away from the person, outcome, or constraint that matters.

1.3.0current version
4linked stages
MITopen source
Purpose is the through-line

Execution can be flawless and still miss the point.

Match the weight of the check to the stakes.

Purpose Check gives an agent a simple choice architecture. It knows when to stay quiet, when to state an assumption, and when one question is worth the interruption.

Gate 01 / low consequence

Check internally. Keep moving.

A clear, low-risk task gets a brief purpose check and no question. The work should feel fast because the stakes are low.

Proceed without asking

Four stages. One anchor.

The purpose anchor carries the original intent from the first decision to the final handoff. It is one sentence naming who this is for and what gets better for them.

01 / SHOULD THIS EXIST?

Start with the person.

Identify who benefits, whether the request is the real problem, and whether a smaller answer would serve the goal better.

02 / WHAT DOES DONE MEAN?

Make the end state visible.

Define an outcome a non-expert could verify. Name the decision or action the output supports.

03 / JUDGMENT AT THE FORKS

Choose for this context.

Compare meaningful choices against the purpose anchor. Surface only tradeoffs that reach the user or change the stakes.

04 / DOES IT HOLD UP?

Return to the reason.

Check for drift, misplaced quality effort, and unresolved consequential assumptions before delivery.

Give the work a north star.

Keep it concrete enough to disagree with. A purpose anchor is a working constraint, not a mission statement.

Who is this for, and what gets better for them?

Purpose anchor / live

A returning user gets back into their account in under ten seconds.

The skill got better by being challenged.

Five models. Four revisions. The review trail stays in the repository because the corrections are part of the work.

Original review

85

Strong concept. Broad trigger, missing gates, weak completion checks.

Version 1.1

96

Better structure, but the first score rewarded section presence over coherence.

Version 1.2

89

Found the producer-consumer bug: Stage 3 needed an artifact Stage 1 did not always create.

Version 1.3

95

Behavioral contracts resolved. The source is ready for any agent harness.

Make the next call a better one.

Read the complete skill, inspect the review trail, and give your agent a judgment layer that knows when to ask and when to move.