Verified project case study · AI research automation
ResearchSwarm
An extension of karpathy/autoresearch that adds task routing, bounded execution, and persistent operational memory.
Contribution
What ELI Labz contributed
ELI Labz extended the upstream project with a Digital Cognitive Labor routing layer that separates software-executable, Human-action, and hybrid tasks, plus documented safety controls and memory support.
Context
The operating problem
Autonomous research loops can execute digital experiments, but they need explicit boundaries when a request crosses into physical or manual work. ResearchSwarm makes that handoff visible rather than pretending every task is executable by software.
Technical approach
How the system was shaped
- Classified natural-language work into text-based, Human-action, and hybrid domains.
- Added explicit opt-in flags before training actions can execute, leaving planning as the safe default.
- Recorded routing and execution events in a SQLite memory store for later review and reuse.
- Kept the upstream experiment loop's constrained edit and evaluation pattern while adding the routing boundary.
Demonstrated outcomes
What can be inspected
- The repository exposes a runnable CLI for classification, planning, execution, and Human handoff generation.
- The project structure includes a task classifier, router, persistent memory, tests, and analysis notebook.
- Public commit history documents ELI Labz updates to the README and changelog for the extension.
Public sources
Claim boundaries
What this case study does not claim
- • ResearchSwarm is a fork of karpathy/autoresearch. The core training loop and inherited experiment claims are not presented as original ELI Labz work.
- • Example benchmark numbers in the README are not treated here as independently verified production outcomes.
Technology
Apply this engineering approach to your deployment.
Describe the workflow, constraints, and target outcome to generate an initial Forward Deployed Engineer engagement brief.
Scope an engagement