AEF

Agent Engineering Foundation

Your agents.
A shared
foundation.

Bring memory, control, and a measured learning loop to the agents you already use.

  • Replayable runs
  • Tool permissions
  • Reviewed lessons

Developer preview · Python 3.11+ · Local-first setup

One connected workflowInteractive 3D

Context. Action. Reflection.

Retrieve → Prompt agent → Reflect → Consolidate → END. Every workflow node uses shared AEFState.
→ Execution route┄ Shared state

Drag to explore 3D · Arrow keys rotate · Home resets

Each workflow node reads AEFState and returns a StateDelta plus its next route.

Illustrative motion · no live agent execution
Adapted from Red Team Mission Orbit · MIT notice

BUILT FOR THE AGENTS YOU ALREADY USE

Claude · Codex · Grok · Cursor · GitHub Copilot

The AEF platform

Every step connected.
Every run inspectable.

Give your agents a consistent way to retrieve context, act, reflect, and retain useful knowledge. Explore the four steps of an AEF persona workflow.

CONTEXT

Bring relevant context into the run.

The memory-backed retriever supplies context through injected services. Retrieved lessons remain fallible evidence; they cannot grant permissions or override the owner's instructions.

(AEFState, Context, Services)
→ (StateDelta, Route)

Generated persona flow. Custom target graphs require explicit wiring and domain tests.

Replay a recorded run

Checkpoint state and resume interrupted work. Replay re-executes declared deterministic nodes against recorded inputs to check their outputs.

Control tool access

Deny-by-default policies, human approval gates and an audit trail. Target graphs must actually route their tools through these services.

Configure each agent

Knowledge, Policies, Tools, Objectives and Evaluation Metrics vary per agent. Shared state and injected services keep the runtime consistent.

Learning, with evidence

Build a learning loop.
Keep your judgment.

Self-learning here means proposing and evaluating changes to instructions and knowledge. It does not mean retraining model weights or granting an agent permission to rewrite itself.

  1. 01

    Observe

    Capture real failures with run IDs, tool results and the conditions that produced them.

  2. 02

    Propose

    Write one bounded correction with a counterexample and an acceptance check.

  3. 03

    Compare

    Test against the incumbent on frozen tasks. Separate held-out evaluation; include placebo controls where appropriate.

  4. 04

    Review

    Record harm, uncertainty and cost. Reject harmful advice. Passing gates escalates to a human.

A LESSON THAT FAILED

More advice is not always better.

In one owner-repository trial, the learned lesson reduced the measured score. A meaningless placebo bullet scored higher than both. This trial does not establish a general effect; it makes the case for testing advice before adopting it.

Historical trial · ADR 0204 · task-specific scores, not a product benchmark. Learning quality remains unproven.

Incumbent
0.67
Learned lesson
0.40
Placebo
0.87

Take the protocol with you

Reusable learning instructions

A reusable instruction template for scoped work, capability checks, reproducible failures and reviewed lessons. Fill in the task-specific fields before use.

Download learning instructions

Built today: rule-based reflection, bounded prompt proposals and a gated loop harness. LLM reflection is available but off by default; existing trials have not demonstrated task improvement. Live evaluation is opt-in. Automatic merging and evolution remain disabled.

Connect your repository

Your repo.
Your agents. AEF.

Start with a new project or the agents already in your repo. One command selects the target. AEF preserves your instructions and guides the integration there.

Local command preview only. Your path is not sent or stored. Use a macOS or Linux absolute path.

/target-repo "/absolute/path/to/your repo"

Run in a coding-agent session opened in your AEF checkout.

1

Set up AEF once

Clone the AEF repository, create its Python environment, then install .[dev]. This is a separate setup step.

python3 -m venv .venv
.venv/bin/python -m pip install -e '.[dev]'
2

Point and preserve

The skill scaffolds configuration and guides mechanical migration. Native Claude/Grok Markdown and eligible Codex TOML personas get graph registrations. The AEF source stays read-only during invocation.

3

Wire and verify

Install a pinned AEF wheel into the target's own environment. Connect real services, tools and evaluations; run doctor and domain tests. A scaffold alone is not a working agent.

Prefer the terminal? Run the mechanical phase directly.
python3 -I -B /absolute/path/to/aef-core/scripts/target_repo.py '/absolute/path/to/your repo'

Offline by default: no model provider or generated scheduled workflows. This command does not perform semantic wiring. Prompt personas need an explicitly configured model provider before execution. Codex also exposes the skill through /skills where supported.

Research-informed engineering

Built with a view
of the bigger picture.

Primary research and engineering references worth understanding. These connections explain design choices; AEF does not claim to implement every method or reproduce their results.

Graph engineering

Durable stateful execution

LangGraph documents checkpointed state and stores for persistent agent workflows. AEF has its own graph kernel, checkpoint/resume and deterministic replay contracts.

Read the LangGraph documentation ↗AEF: independent implementation

Reflection + memory

Reflexion

The paper explores verbal feedback retained across attempts. AEF offers reflection and memory components, but their presence is not evidence of improved task performance.

Read the paper ↗AEF: related components; gains unproven

Iterative feedback

Self-Refine

Generate, critique, revise: a useful pattern for iterative output refinement. AEF's optional LLM reflection is a related mechanism, not a validated reproduction of the paper.

Read the paper ↗AEF: optional LLM reflection, off by default

Reflective optimization

GEPA

GEPA reflects on execution trajectories and searches prompt candidates using a Pareto frontier. AEF's bounded proposer is simpler. GEPA is a research reference, not an included optimizer.

Read the paper ↗AEF: GEPA not implemented

References reviewed September 12, 2026. No cross-framework ranking or third-party benchmark result is claimed for AEF.

Product availability

Clear about what ships.

Implemented Graph execution, state, checkpoint/replay, policy and audit, memory, context retrieval, evaluation, OTel tracing, rule-based reflection and gated candidate review.

Requires target wiring Domain behavior, actual tool containment, provider configuration, durable services and meaningful evaluation.

Unavailable or disabled Multi-agent coordination is stubbed; fan-out is not executed. There is no knowledge-graph service or general planner. Evolution and automatic merging are disabled.

Meet your next workflow

Bring it all together
with AEF.

Start with your repository ↗