# Agent systems are learning what to preserve, not only what to retrieve.

Five source-checked changes in orchestration, execution evidence, memory validity, MCP discovery, and local inference that affect implementation choices.

Published: 2026-10-09 | Updated: 2026-10-09
Author: /init editorial
Method: AI-assisted, source-checked
Canonical: https://init.news/notes/agent-systems-brief-2026-10-09/

## The answer

The strongest change this week is architectural: agent systems are beginning to preserve coordination evidence as first-class state. That means successful handoffs, execution receipts, memory provenance, policy boundaries, and runtime migrations need explicit representations. The practical move is not to add another autonomous agent. It is to decide which state must survive a run, which state may be reused, and which state must be invalidated or withheld.

## The five changes worth carrying into design review

Rank and source | What changed | Decision affected | Read this
--- | --- | --- | ---
1. WorkflowOps — preprint, 6 Oct | The orchestrator keeps a transition matrix of successful agent handoffs and uses it as soft guidance when constructing a dependency-valid DAG. Embeddings handle routine capability matching; an LLM is reserved for uncertain cases. [1] | Persist successful workflow structure separately from conversational memory. Test a cheap, inspectable prior before training a router or regenerating every workflow from zero. | Sections 4.1–4.3 for matching, priors, and agent creation; Section 5 for experiments; Section 6 for limits; Appendices G–I for cards, messages, and configuration.
2. Beyond Corrected Memory — preprint, 6 Oct | The paper argues that correct shared records do not prove that an execution complied with its duties. It evaluates explicit evidence for state use, handoff receipt, action dependence, and final-state agreement. [2] | Design traces to answer whether an agent received, used, and acted on required information. A message log alone is not an execution audit. | The execution-consistency definition, controlled evidence-removal study, CAVERT diagnostic pipeline, and benchmark/executor results.
3. PACMI — preprint, 5 Oct | Memory updates become a provenance problem: new evidence can supersede one record while forcing dependent claims into different validity states. Retrieval then considers validity and checks whether a query assumes a stale premise. [3] | If a memory graph informs decisions, store typed dependency edges and validity state. Do not treat deletion, recency, or embedding similarity as sufficient invalidation. | Figure 1, Section 4, the node/context/answer evaluation, and the stated assumptions on annotated graphs and two-hop diagnostic cases.
4. Uber MCP Gateway — production report, 1 Oct carry-in | Uber separates an MCP registry control plane from a proxy data plane, auto-discovers interfaces into a disabled state, and exposes progressive discovery, response projection, and a code-oriented client instead of loading every schema into context. [4] | Once the tool estate is larger than a curated prompt, centralize ownership, policy, discovery, and observability. Give agents a narrow discovery path instead of the entire catalog. | AutoCrawler and authoring controls; the data-plane diagram; Runtime Discovery, Response Projection, and Code Mode.
5. Ollama 0.40.2 — release, 8 Oct | Previously downloaded models are upgraded in the background on first use for llama.cpp compatibility. Ollama temporarily keeps the original data so a downgrade remains possible. [5] | Treat a local-runtime upgrade as a storage migration: budget temporary disk duplication, pin the rollout, and test rollback or re-pull behavior before updating every node. | The Model upgrades note and its backup-removal command; this is release behavior, not a published performance benchmark.

## What actually changed in the architecture

These sources converge on a boundary that older agent diagrams often hide. Conversation state is not enough. An operational system also needs orchestration history, causal execution evidence, memory lineage and validity, tool ownership and policy, and runtime artifact state. Each is updated on a different schedule and should have a different retention rule.

- Keep task dependencies hard and learned collaboration history soft. A prior may suggest an edge or agent, but it should not override the task's required ordering.
- Record receipts and action dependence at execution time. Reconstructing them from prose logs later is weaker and more expensive.
- Make memory writes and invalidations policy-aware. Relevant information can still be stale, historically valid, or inadmissible for the current principal.
- Separate MCP discovery from exposure. Generate or discover tool definitions into a reviewable registry, then enable them deliberately with authorization and response controls.
- Version local-model artifacts and their runners together. A compatible migration path is part of the runtime contract, not housekeeping.

## Evidence limits

WorkflowOps, Beyond Corrected Memory, and PACMI are author-reported preprints; this review did not reproduce their experiments. WorkflowOps bootstraps roughly 65 synthetic workflows for a 35-agent matrix and reports its strongest gains on structured tasks where handoffs transfer. PACMI supplies the dependency graph in its diagnostic benchmark, and its cascading propagation did not yield a statistically significant final-answer difference at the evaluated scale. Uber's scale and benefits are a first-hand production report, not an independent audit. The Ollama item is directly observed release behavior, not evidence that every model becomes faster.

## Review window and selection rule

The primary review window was 2–9 October 2026. The 1 October Uber report is included as a one-day carry-in because this is the first agent-systems edition and its control-plane pattern changes a concrete implementation decision. Routine model additions, papers without an inspectable decision boundary, and generic multi-agent surveys were not ranked. No fresh GraphRAG result cleared the bar more strongly than the memory-validity work above.

## Sources

- [WorkflowOps: Learning Agent Collaboration Priors for Multi-Agent Workflow Orchestration](https://arxiv.org/abs/2610.07860) — [1] Submission date, method, author-reported experiments, implementation appendices, and limitations.; checked 2026-10-09
- [Beyond Corrected Memory: Execution Consistency in Multi-Agent Systems](https://arxiv.org/abs/2610.08101) — [2] Submission date, execution-consistency definition, controlled evidence study, and author-reported CAVERT results.; checked 2026-10-09
- [PACMI: Provenance-Aware Cascading Memory Invalidation for Long-Term LLM Agents](https://arxiv.org/abs/2610.05732) — [3] Provenance graph, validity lattice, retrieval and premise checks, benchmark scope, results, and limitations.; checked 2026-10-09
- [Uber Engineering: Designing MCP Gateway, Uber's MCP Management Platform](https://www.uber.com/ca/en/blog/designing-mcp-gateway/) — [4] Publication date, reported production architecture, scale, discovery, authorization, response projection, and code-mode design.; checked 2026-10-09
- [Ollama v0.40.2 release notes](https://github.com/ollama/ollama/releases/tag/v0.40.2) — [5] Release date, background model migration, temporary backups, cleanup, and downgrade behavior.; checked 2026-10-09
