PASS
Consolidate verified output and update the blueprint before the next task. Source acceptance and production release are separate decisions.
Method
An Implementation-Planning-Driven AI-native engineering method: from customer problems, blueprints and task cycles to two verification tracks, handover and production adoption.
There is no optimal decision, only sustainable ones. Rule out known bad paths with tolerance and red lines, and keep the choice in play. Decompose a problem into milestones with dependencies before touching code.
Feature lists take the technical view; business event flows take the business view. Event Storming asks "what events happened", producing event flows, bounded contexts, and state machines. Every rejected option lands in a divergence record: decision, reason, scope, date. The output is REQUIREMENTS.md.
A prototype answers questions; it is not the deliverable. When requirements are doubtful, interactions complex, or stakeholders need convincing, prototype first. Three fidelities: text validates logic, wireframe validates structure, high-fidelity validates interaction. The mock-to-real evolution lands in POC-MANIFEST.md, each piece with an honest ledger.
A blueprint is not documentation; it is a constraint device. CONTEXT.md records facts, ARCHITECTURE.md defines constraints, AGENTS.md governs reading and execution, and CHANGELOG.md records changes. Negative-space design: fence off the paths not to take, then let AI generate inside what remains.
Decompose, dispatch instructions, code, inspect, decide the branch, update the blueprint. Milestones are engineering-dimensioned, not requirement-dimensioned: database, API, validation, and frontend each land as separate segments, independently inspected. The verdict has exactly three values: PASS, NEEDS_FIX, REBUILD.
Consolidate verified output and update the blueprint before the next task. Source acceptance and production release are separate decisions.
Record defects, impact and repair criteria; rerun failed categories and affected regression. Reassess structure and cost when failures repeat.
Rebuild after reviewing structure and cost. Preserve work and verify target commits and branch sharing. Do not rewrite shared history; recover deployment and external effects separately.
No work without a blueprint. No consolidation without inspection. Rebuild on chaos. Each discipline carries the cost of breaking it — use the standards to reduce risk and validate their effect. Preserve work and check targets and shared-history boundaries before rebuilding.
No code before requirements, architecture, and frontend design produce their docs. Otherwise drift starts at step one.
No commit or consolidation before inspection. Otherwise drift compounds.
Stop, save verified facts and preserve work before restoring the persisted blueprint. Compare rebuilding costs and check targets and shared-history boundaries.
Confirm the plan and task authorization, then execute within scope. Escalate changed requirements, permissions or customer commitments. Sanitize telemetry and retain runtime records; update the blueprint only with verified facts.
Stop when APIs are invented, constraints are lost, failures repeat or fixes spread errors. Save verified facts and current work before using the tool-supported reset. Restore the blueprint, current task and acceptance criteria, then restate and verify them. Window usage is a diagnostic signal rather than a universal threshold.
The chain makes demands of its practitioner; the practitioner grows into them. Eight capabilities: problem decomposition, end-to-end delivery, reverse understanding, fast validation, two-way translation, honest ledger, context self-care, self-executing discipline.
Turn a vague ask into milestones with dependencies. Blueprint first, then code.
Own scoping, engineering, handover and adoption, with clear responsibilities and escalation.
Read the structure of an unfamiliar system from its live site and legacy code. Without being taught.
Use a prototype to test key assumptions, with timing based on complexity. Prototypes answer questions; they are not the deliverable.
A blueprint is a technical document the business can read, and a business document the engineers can read.
Know what you do not know. Three-level fidelity marks; never present inference as fact.
Know when to reset. Rebuild before context rots; never push through.
Blueprint checks, tests and evaluations provide inspectable evidence; people retain review and release responsibility.
Software tests verify code contracts; evaluations verify probabilistic outputs. Golden sets need representative cases, held-out samples and versions; judges require human calibration. Check prohibitions before scores, including orchestration changes affecting inputs or calls.
Customer projects follow scoping, validation and delivery, with scope and ownership explicit. Customer owners rerun evaluations and recovery. Observe real-task adoption over a complete 30-day window separately from business impact. Feed field knowledge back within authorized boundaries.
Engineering
The toolchain of the method. 22 installable skills with quality governance and a test grid.
22 total / 14 Engineering / 3 Product / 5 RC Philosophy
See the project README for actual repository checks. This section explains skill methods and does not claim measured model evaluation. Evaluation →
Skills organize implementation without granting customer permissions.
Scope and blueprint first. Job builds within authorization; owners decide release and customer commitments.
Use Workflow for scoped tasks, Coach for guidance and Orchestrator for dependencies.
Inspector checks the blueprint; QA verifies software. Evaluate changed model inputs or call paths separately.
Advisor provides recommendations; people own assumptions and decisions.
POC tests assumptions, Cloner reconstructs references, Legacy Recon maps old systems, Next resumes tasks.
Review changes against authorization, contracts and responsibilities.
Core contract changes: check authorization, impact and severity. Stop propagation, preserve work, then choose repair or rebuild.
Needless dependencies and abstractions: review against real task needs and keep simpler module boundaries.
Growing size and coupling: use file length as a review signal and split by responsibility, avoiding mechanical line-count verdicts.
Retain separate software and model verification evidence.
Software track: unit checks for rules, integration for interfaces and critical-path checks for generated pages and user flows. Record scope and results.
Model track: golden sets, human-calibrated judges and versioned regression. Changed inputs or orchestration also trigger evaluation. This site has no production model calls.
Prohibitions first: do not lower gates to pass failures. Rerun failed categories and affected regression after repair, retaining item evidence and recovery checks.
Evaluation
Running code and reliable model behavior require distinct judgments. Representative cases, calibrated judges and comparable regression provide reproducible release evidence.
This section explains textbook methods. This static content site has no production model calls or measured model evaluation results.
01
Build deduplicated cases from real tasks, authorized failures and boundaries. Record provenance, input snapshots, expected actions, prohibitions and categories. Smoke is a regression subset; report adversarial cases separately without double counting. Keep tuning and held-out samples separate. Sample count does not establish readiness, and risk-weighted pass rates are not production accuracy.
02
Use deterministic checks for labels, quantities, fields and format where possible; an LLM judge can assist semantic review. Define observable behavior for each rubric grade. Have humans independently label and arbitrate disagreements. Report agreement separately for routine and dangerous boundaries; check position, self-preference and length biases. Validate judge updates on held-out samples; aggregate agreement alone is insufficient.
03
Record dataset, input snapshot, model, prompt, rules, judge, orchestration and runtime configuration. Input or call-path changes also require evaluation. Rerun candidate and baseline under matching conditions, re-evaluating both after dataset or judge updates. Retain item outputs and category statistics. Rerun failed categories and affected regression after fixes; version newly added cases.
04
Check prohibitions and severity before correctness, unsupported claims, format, cost and latency, reviewing regression by category. Agree on gates before changes. Lower cost or unchanged aggregate scores cannot excuse unauthorized actions or fabrication. Stop candidate release on failure; preserve work, restore verified configuration and recheck. Source, deployment and external effects need separate recovery plans.
Do types, rules, interfaces and critical user paths meet contracts? Record tests, build evidence, the verified version and failures.
Check correctness, unsupported claims, format, cost and latency on representative inputs. Calibrate judges with humans. This site currently has no production model calls.
During a complete 30-day post-release window: distinct target users completing real tasks / the predefined target cohort. Specify membership, eligibility, access, timezone and events. Login and training do not count as use.
Compare time, rework and errors at the same workflow stage. Explain confounders such as seasonality, training and workload, including review, deployment and maintenance costs. Adoption growth alone does not establish value.
Based on textbook chapters 21–23, section 26.3 and chapter 39. Dataset size and example gates must be agreed for each project risk. Curriculum →
Further reading: Anthropic · Demystifying evals for AI agents · Judging LLM-as-a-Judge
Delivery
AFDE owns verifiable outcomes in customer production. Scope the task, validate with evidence, then establish handover and sustained adoption.
Observe frontline events. Confirm the problem, data sources, quality and access. Record the baseline, initial scope, prohibitions, owners and escalation. A synthetic-data demo cannot establish production readiness.
Exit: draft requirements, data inventory, seed cases and validation entry criteria. Reduce scope or stop if evidence is insufficient. Authorized customer owners confirm scope and commitments.
Build representative evaluation cases in an authorized environment, with expected actions and prohibitions. Freeze the baseline, versions and judge. Compare candidates under identical conditions, retaining item outputs and failure categories. Verify software and model behavior separately.
Exit: reproducible evaluation, known failures and mitigations, and a proceed or stop decision. Aggregate scores, cost and latency gains cannot override prohibited behavior.
Rehearse configuration recovery, expired permissions and interface failures before shadow operation and a controlled pilot. Put evidence into the existing workflow, with human confirmation, rejection and fallback. Field observation complements asynchronous records.
Exit: customer owners independently rerun evaluations, explain metrics and rehearse recovery. Business and technical owners sign off. Deployment, 30-day adoption and business impact each require their own evidence.
Record an observation, applicability and evidence version, separating unverified hypotheses. Share only in authorized spaces; raw customer data stays outside team assets.
Describe the problem, impact, reproducible evidence, boundaries and next step. Name the recipient, maintainer and review date. Validate changes in the field.
Create playbooks, skills or tools with synthetic reproductions, applicability, versions, owners and failure boundaries. Calibrate through sharing and practice.