Replay a recorded run
Checkpoint state and resume interrupted work. Replay re-executes declared deterministic nodes against recorded inputs to check their outputs.
AEF
Agent Engineering Foundation
Bring memory, control, and a measured learning loop to the agents you already use.
Developer preview · Python 3.11+ · Local-first setup
Drag to explore 3D · Arrow keys rotate · Home resets
Each workflow node reads AEFState and returns a StateDelta plus its next route.
Illustrative motion · no live agent execution
Adapted from Red Team Mission Orbit · MIT notice
Claude · Codex · Grok · Cursor · GitHub Copilot
The AEF platform
Give your agents a consistent way to retrieve context, act, reflect, and retain useful knowledge. Explore the four steps of an AEF persona workflow.
The memory-backed retriever supplies context through injected services. Retrieved lessons remain fallible evidence; they cannot grant permissions or override the owner's instructions.
(AEFState, Context, Services)
→ (StateDelta, Route)Generated persona flow. Custom target graphs require explicit wiring and domain tests.
Checkpoint state and resume interrupted work. Replay re-executes declared deterministic nodes against recorded inputs to check their outputs.
Deny-by-default policies, human approval gates and an audit trail. Target graphs must actually route their tools through these services.
Knowledge, Policies, Tools, Objectives and Evaluation Metrics vary per agent. Shared state and injected services keep the runtime consistent.
Learning, with evidence
Self-learning here means proposing and evaluating changes to instructions and knowledge. It does not mean retraining model weights or granting an agent permission to rewrite itself.
Capture real failures with run IDs, tool results and the conditions that produced them.
Write one bounded correction with a counterexample and an acceptance check.
Test against the incumbent on frozen tasks. Separate held-out evaluation; include placebo controls where appropriate.
Record harm, uncertainty and cost. Reject harmful advice. Passing gates escalates to a human.
In one owner-repository trial, the learned lesson reduced the measured score. A meaningless placebo bullet scored higher than both. This trial does not establish a general effect; it makes the case for testing advice before adopting it.
Historical trial · ADR 0204 · task-specific scores, not a product benchmark. Learning quality remains unproven.
Take the protocol with you
A reusable instruction template for scoped work, capability checks, reproducible failures and reviewed lessons. Fill in the task-specific fields before use.
Built today: rule-based reflection, bounded prompt proposals and a gated loop harness. LLM reflection is available but off by default; existing trials have not demonstrated task improvement. Live evaluation is opt-in. Automatic merging and evolution remain disabled.
Connect your repository
Start with a new project or the agents already in your repo. One command selects the target. AEF preserves your instructions and guides the integration there.
Local command preview only. Your path is not sent or stored. Use a macOS or Linux absolute path.
/target-repo "/absolute/path/to/your repo"Run in a coding-agent session opened in your AEF checkout.
Clone the AEF repository, create its Python environment, then install .[dev]. This is a separate setup step.
python3 -m venv .venv
.venv/bin/python -m pip install -e '.[dev]'The skill scaffolds configuration and guides mechanical migration. Native Claude/Grok Markdown and eligible Codex TOML personas get graph registrations. The AEF source stays read-only during invocation.
Install a pinned AEF wheel into the target's own environment. Connect real services, tools and evaluations; run doctor and domain tests. A scaffold alone is not a working agent.
python3 -I -B /absolute/path/to/aef-core/scripts/target_repo.py '/absolute/path/to/your repo'Offline by default: no model provider or generated scheduled workflows. This command does not perform semantic wiring. Prompt personas need an explicitly configured model provider before execution. Codex also exposes the skill through /skills where supported.
Research-informed engineering
Primary research and engineering references worth understanding. These connections explain design choices; AEF does not claim to implement every method or reproduce their results.
Graph engineering
LangGraph documents checkpointed state and stores for persistent agent workflows. AEF has its own graph kernel, checkpoint/resume and deterministic replay contracts.
Read the LangGraph documentation ↗AEF: independent implementationReflection + memory
The paper explores verbal feedback retained across attempts. AEF offers reflection and memory components, but their presence is not evidence of improved task performance.
Read the paper ↗AEF: related components; gains unprovenIterative feedback
Generate, critique, revise: a useful pattern for iterative output refinement. AEF's optional LLM reflection is a related mechanism, not a validated reproduction of the paper.
Read the paper ↗AEF: optional LLM reflection, off by defaultReflective optimization
GEPA reflects on execution trajectories and searches prompt candidates using a Pareto frontier. AEF's bounded proposer is simpler. GEPA is a research reference, not an included optimizer.
Read the paper ↗AEF: GEPA not implementedReferences reviewed September 12, 2026. No cross-framework ranking or third-party benchmark result is claimed for AEF.
Product availability
Implemented Graph execution, state, checkpoint/replay, policy and audit, memory, context retrieval, evaluation, OTel tracing, rule-based reflection and gated candidate review.
Requires target wiring Domain behavior, actual tool containment, provider configuration, durable services and meaningful evaluation.
Unavailable or disabled Multi-agent coordination is stubbed; fan-out is not executed. There is no knowledge-graph service or general planner. Evolution and automatic merging are disabled.
Meet your next workflow