Awesome Artificial Intelligence
Section: Software factories and agent orchestration · OpenAI's field report on building software with coding agents, repository constraints, automated checks, and human steering.
Entry
Appears in 4 awesome lists
OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.
Section: Software factories and agent orchestration · OpenAI's field report on building software with coding agents, repository constraints, automated checks, and human steering.
Section: Foundations · OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.
Section: Harness Engineering · Official OpenAI post: "leveraging Codex in an agent-first world"
Section: Related Resources · Environment design, intent, feedback loops, repo-as-system-of-record
Affaan Momin's agent-harness operating system: 68 specialized agents, 286 skills, hooks, memory, continuous learning, and AgentShield security scanning across Claude Code, Codex, Cursor, OpenCode, and other harnesses. The clearest open-source example of packaging an end-to-end engineering workflow…
LangChain's batteries-included agent harness (released April 2026) with built-in planning, filesystem tools, shell access, sub-agents, and auto-summarization. The clearest open-source demonstration of how a general-purpose coding agent harness can be made ready-to-run out of the box while…
Anthropic's pattern for maintaining agent progress across multiple context windows: an initializer agent sets up the environment once and hands off to a coding agent that makes incremental progress each session. The structured handoff mechanism — feature lists, git commits, and test gates as…
Swaps the filesystem and bash providers for a mirage virtual workspace: file tools and shell commands run over mounted resources (RAM, S3, Redis, Slack, Gmail, Notion, Postgres) instead of the host disk, with per-mount read/write/exec modes, per-command sandbox routing (monty, pyodide, quickjs in…
"Stop prompting. Design the loop." — practical patterns, starters & CLI (loop-audit, loop-init, loop-cost) for systems that discover work, hand it to agents, verify results, and persist state across Claude Code, Codex, Grok, and OpenCode; report-only week one, scores loops on a "Loop Ready" rubric…
Anthropic's framework for evaluating agent behavior: what to measure, how to build eval harnesses, and why unit-test-style evals fail for agents.
by Anthropic - Anthropic's foundational taxonomy of agent patterns — prompt chaining, routing, orchestrator-workers, and evaluator-optimizer — and when to use each.
by Anthropic - A practical account of orchestrator and subagent coordination, prompt design, and evaluation that maps directly to Claude Code's subagents and agent teams.