Skip to content

The Feedback Loop Is the Agent

#llm-agents #agent-benchmarks #domain-automation #optimization #agent-memory

For the last year I've watched agent benchmarks quietly change what they measure. The old question was "can the model answer this?" The new question is "can the agent finish the job?" That shift sounds small. It isn't. Answering takes one forward pass. Finishing takes a loop: act, check, diagnose, retry, remember.

Five papers landed in the same August 2026 arXiv window, each attacking a different domain. FormuEvo generates mixed-integer programming formulations. NetConfArena evaluates agents that configure emulated multi-device networks. VideoRover unifies video reasoning with deep research. DG-Mem bolts external memory onto frozen multimodal models. LongWoF-Bench asks whether verified execution experience can be reused across runs. On GitHub, the awesome-llm-apps repo keeps growing past 100 open-source agents, which tells you what practitioners ship when they stop waiting for benchmarks.

The thread through all of it: the model is the cheap part. The loop around it is where the work happens.

FormuEvo: when the solver is the judge ​

Mixed-integer programming sits at the core of operations research. Airlines use it for crew scheduling. Logistics companies use it for routing. The standard workflow is painful: a human expert translates a business problem into a MIP formulation, then hands it to a solver like Gurobi or SCIP, then waits.

LLMs can now generate MIP models from natural language. But they optimize for semantic correctness, not formulation strength. A semantically correct formulation can be computationally hopeless. The same problem, formulated differently, can solve 5.5x faster. That gap is the difference between an hour of solver time and about eleven minutes.

FormuEvo treats formulation design as evolutionary optimization over executable modeling programs. An LLM generates candidates, a solver evaluates them, and the LLM performs crossover, mutation, and repair based on what the solver reports. The key trick is the diagnosis mechanism: instead of blind mutation, FormuEvo reads fine-grained solver statistics as verbal gradients that guide the next generation. Things like node counts, bound gaps, and cut violations become instructions the model can act on.

Key numbers: 5.5x solver speedup over expert-designed formulations. Hour-long solves drop to about 11 minutes. Knowledge transfers to unseen problems and smaller LLMs.

The transfer result matters more than the speedup. FormuEvo distills what it learns into reusable modeling strategies, then bootstraps smaller models with them. That's the same pattern we'll see in three more papers: capture the experience, externalize it, reuse it.

NetConfArena: benchmarks that actually execute ​

Network configuration is where I'd expect agent automation to fail first. Configuring a router means reasoning about protocol interactions, topology dependencies, and vendor quirks. A config that parses correctly can still break the network.

Most existing benchmarks treat configuration as static command generation. You compare the agent's output to a reference string. NetConfArena argues that's the wrong abstraction, and I think it's right. The paper places agents in emulated multi-device networks, gives them a compact action interface, and evaluates the resulting network behavior with hidden executable test cases. The hidden tests check whether the network works afterward. Command text is never compared.

The scale matters: 480 task instances from 96 protocol-focused templates, producing 3,840 execution trajectories. That's enough data to see failure patterns, and the patterns are not what you'd expect. Failures are not mostly command errors. They're gaps in task-specification adherence and planning. The agent understood the protocol. It just didn't follow the full spec or plan across devices.

This is the same lesson as FormuEvo, applied to evaluation. You can't improve what you can't execute. A benchmark that only checks syntax will tell you your agent is fine when it isn't.

VideoRover: tool use as perception ​

Video reasoning has a specific failure mode: the evidence you need is often not in the video. A question about a historical event shown in archival footage requires external knowledge the model doesn't have and can't see in the frames. Thinking-with-Videos handles temporal perception. Deep Research handles multi-step search. Nobody had put them together.

VideoRover does. Given a video-question pair, it iteratively coordinates three tools: video cropping, multimodal search, and webpage browsing. Each tool result selects the next action. A localized clip suggests a search query. A search result triggers closer video inspection. The tools act as a perception system, not a bolt-on.

The training data comes from an automated curation pipeline: 26K verified SFT trajectories and 3K reinforcement learning instances. The payoff is that VideoRover-8B-RL matches proprietary models when answering directly, no tools involved, and beats larger open-source models when both sides get the same tool suite.

An 8B model matching proprietary models matters because of what it implies about deployment. That's a model that fits on a single GPU, no cloud API required. The capability comes from the loop: the coordination between perception, retrieval, and long-horizon RL. Scale is secondary.

Quick Take: In every system here, the gains come from the loop around the model, not from the model itself.

Two papers, one answer on memory ​

DG-Mem and LongWoF-Bench approach the same problem from opposite ends. Both ask what to do with experience after a successful run.

DG-Mem works on multimodal scientific reasoning. It augments a frozen MLLM with non-parametric, externally stored memory built from training-time rollouts. The design follows Complementary Learning Systems theory from human memory research: an instance-grounded exemplar memory plus a category-level schema memory of IF-THEN rules. A transient reflection store sits between them, so schemas are synthesized only from abstract reflections, never from raw exemplar text.

Two details stand out. First, an online concept categorizer grows the category space during training instead of committing to a fixed taxonomy. Second, a Shapley context attribution procedure decomposes correctness across the retrieved rule set, producing per-rule utility weights that re-weight retrieval at test time. No gradient updates anywhere. That means it works on closed-weight models and on-device backbones. DG-Mem improves results consistently across four backbones: Qwen3.5-27B, Qwen3.5-122B-A10B, GPT-5-Nano, and Gemini-3-Flash.

LongWoF-Bench asks the same question for long workflows. The benchmark contains 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved Gene outperform Skill by 8.7 to 15.5 percentage points across all seven evaluated models, and the gains extend to consumer models from different families.

Reference-distilled Gene don't show the same advantage. Compact representation alone isn't enough. The utility comes from verified experience provenance. You can't compress your way to good behavior; you have to earn the experience through verification first.

For Claude Opus specifically, Gene reuse completed 39 more tasks than Skill while cutting solve-time token consumption by 9.9%. Fewer tokens because the model doesn't rediscover failure modes from scratch every run.

SystemDomainCore mechanismResultPractical meaning
FormuEvoMIP formulationLLM-guided evolution, solver diagnosisUp to 5.5x solver speedupHour-long solves become ~11 minutes
NetConfArenaNetwork configEmulated networks, hidden tests3,840 trajectories analyzedStatic benchmarks hide planning failures
VideoRoverVideo researchCropping + search + browsing loop8B model matches proprietarySingle-GPU deployment, no cloud API
DG-MemMultimodal reasoningExternal exemplar + schema memoryGains on 4 open and closed backbonesMemory without fine-tuning
LongWoF-BenchLong workflowsEvoMap Gene from confirmed runs+8.7 to +15.5 pts across 7 modelsExperience transfers across model families

The shared loop ​

Every system in this cluster runs the same loop. This diagram is the pattern all five papers converge on.

FormuEvo's executable check is the solver. NetConfArena's is the hidden test suite. VideoRover's is the next tool result. DG-Mem and LongWoF-Bench add the memory node on the right. The differences are in the domain, not the architecture.

The failure node matters as much as the success path. Every paper that reports failure analysis finds the same thing: agents don't fail because they can't generate text. They fail because they can't diagnose their own mistakes, or they don't follow the full specification, or they don't remember what worked last time.

What 100+ open-source agents actually teach you ​

I cloned awesome-llm-apps on a Tuesday afternoon expecting a pile of toy demos. It's not that. The repo is 100+ hand-built agents, all Apache-2.0, and it's structured like a taxonomy of what works in production.

The first thing I noticed: the agents that feel solid are the ones with tight output contracts. The chess agent validates moves. The trust-gated research team puts every action in a hash-chained audit trail. The typed RAG agent returns exact citations or refuses to answer. The flimsier ones, the ones that feel like demos, are the ones where the output is just prose with no check attached.

I ran the travel agent first because it's the canonical starter. Single file, one API key, streamlit run travel_agent.py, done. It worked. Then I tried the project graveyard skill, which is aimed at coding agents: it finds abandoned side projects in your repo history and tells you why each one died. That one's clever, because it has a verifiable input (git history) and a useful output (a diagnosis you can check against your own memory).

The npx skills add flow shows how far this has come. One command installs a skill into Claude Code, Codex, or Cursor, and every skill ships real code that passes a security and eval CI gate. That's the feedback loop idea applied to the agent development process itself. The skills that survive are the ones with checks.

The repo also confirms where the field's energy is going. Production-style agents are the biggest category at 22 templates. RAG pipelines follow at 21, ranging from basic chains to corrective RAG that grades its own retrieval. The agent teams section shows orchestration maturing: routing, fallback, handoffs. And the token optimization section, with claims of 30-90% cost reduction, shows where the money goes.

My take after a week with it: the repo is a better guide to agent design than most papers, because it shows you the distribution of what people actually ship. The long tail is narrow. Most working agents are either single-file tools with one API key or small teams with explicit handoffs. The middle ground, the sprawling multi-agent system with no verifier, is where projects go to die.

What trips people up ​

I've now read five papers and one repo's worth of agent failures. The same mistakes keep showing up.

The first is confusing syntax with correctness. NetConfArena's whole premise is that a config which parses can still break the network. The same applies everywhere: a well-formed output is not a correct one. If you don't have an executable check, you don't have a signal.

The second is ignoring formulation strength. FormuEvo's 5.5x speedup comes entirely from reformulation, not from a better solver or a bigger model. If you're using LLMs to generate code or models for downstream tools, semantic correctness is the floor, not the goal. The solver's internal statistics are the signal that matters.

The third is throwing away verified trajectories. LongWoF-Bench shows that reference-distilled experience doesn't transfer, but verifier-confirmed experience does. If you're logging agent runs, log the ones that passed checks separately and structure them for reuse. Raw conversation logs are not memory.

The fourth is memory as a text dump. DG-Mem's schema memory works because schemas are synthesized from abstract reflections, not pasted exemplar text, and because Shapley attribution re-weights retrieval per rule. Sticking transcripts in a vector store and calling it memory is the equivalent of a filing cabinet with no index.

The fifth is having no diagnosis step. The papers that report failure analysis all find that agents struggle to recover from their own mistakes. A retry loop that just re-samples the same prompt is not a feedback loop. You need structured diagnosis: solver statistics, test failures, tool errors, converted into something the model can act on.

One thing to remember: the difference between a demo agent and a production agent is the presence of a verifier. Everything else, the memory, the diagnosis, the retry logic, is built on top of that.

The Bottom Line ​

If you're building an agent that produces artifacts for downstream tools, adopt a solver-informed or test-informed feedback loop like FormuEvo's. Semantic correctness is not enough; your agent needs to see the downstream tool's internal statistics to improve.

If you're evaluating agents in a domain with real-world consequences, skip static benchmarks and build an executable environment like NetConfArena's. Hidden test cases on emulated infrastructure will tell you more in a week than string-matching evaluations will tell you in a year.

One thing to watch: verified experience reuse is moving fast. LongWoF-Bench and DG-Mem both show that externalized, verifier-confirmed memory transfers across models and shrinks token costs. Within six months, expect this pattern to be a standard feature in agent frameworks, not a research novelty.