Multi-Agent Systems · Materials Discovery

NeverOT

Agent-Native Orchestration for
Never-Ending Autonomous Optimization

Sissi Feng · Brandon R. Sutherland · Ali Shayesteh Zeraati · Yang Bai
Acceleration Consortium · University of Toronto
Multi-Agent LLM Orchestration Self-Driving Lab AI for Science
Motivation

Current Lab Automation Is Brittle

⚙️ Primitive Workflow Approach

  • Every step manually specified
  • Assumes deterministic agent behavior
  • Breaks on edge cases, requires human recovery
  • Combinatorial explosion as tasks scale
  • "Don't trust agents — control everything"

🤖 Agent-Native Approach

  • Goals specified, not steps
  • Agents handle local uncertainty
  • Graceful degradation and recovery
  • Hierarchical abstraction scales naturally
  • "Trust agents where they are reliable"

"Represent — and control — a system at the level at which it can truly be known."

— Feng et al., Matter 9, 102652 (2026) · Purposeful Imperfection
Vision

What is NeverOT?

A never-ending autonomous optimization framework where a hierarchy of LLM-powered agents continuously closes the loop between hypothesis, experiment, and discovery — without human intervention in the inner loop.

🎯

Goal-Driven

High-level scientific objectives decompose automatically into actionable sub-tasks across agent layers.

🔄

Closed-Loop

Experiment results feed back into hypothesis generation, creating a continuous learning cycle with no dead ends.

🧠

Multi-Agent

Specialized agents at each abstraction layer — reasoning, planning, execution — collaborate with well-defined interfaces.

📐

Fidelity-Matched

Control granularity adapts to agent reliability — coarse where uncertain, fine where mechanisms are stable.

Architecture

Three-Layer Agent Hierarchy

Each layer operates at its own abstraction level, only passing goal states — never implementation details — to the layer below.

L3 · InverseDesignAgent
Hypothesis generation · Goal inversion · Strategy selection · Long-horizon planning
↓ target property space ↓
L2 · DesignAgent
Experimental design · Parameter optimization · Feasibility checking · Batch scheduling
↓ experiment parameters ↓
L0 · Hardware Interface
Instrument control · Data acquisition · Error handling · Result streaming
↓ observations & results ↑

Key Design Principles

  • 🔒 Interface contracts between layers — no leaking implementation
  • 🔁 Async feedback — L3 keeps planning while L0 executes
  • ⚖️ Uncertainty propagation — failure modes bubble up, not crash
  • 🧩 Pluggable backends — swap hardware without touching reasoning
  • 📊 22 unit tests passed on first version
Multi-Agent Collaboration

How Agents Talk to Each Other

🧠
InverseDesignAgent (L3)
Coordinator · Long-horizon planner · Hypothesis owner
⬇   ⬆
⚗️
DesignAgent A
Synthesis pathway
📊
DesignAgent B
Characterization
🔬
DesignAgent C
Activity screening
⬇   ⬆
🤖
Hardware Layer (L0)
Liquid handler · Plate reader · Potentiostat · Data logger

Message Passing

L3 → L2: "Maximize OER activity in Fe-Ni-Co space, overpotential < 300 mV"
L2 → L0: { composition: [Fe:0.4, Ni:0.4, Co:0.2], pH: 13, scan_rate: 5mV/s }
L0 → L2: { overpotential: 287mV, stability: 2.1h, status: ok }
L2 → L3: "Candidate #47 exceeds threshold. Confidence: 0.87. Suggesting refinement in Fe:0.35–0.45."
Core Component

InverseDesignAgent — L3

The reasoning core: given a target property, it works backward to propose which compositions and conditions are most likely to achieve it — and why.

Capabilities

  • 🔭 Goal inversion — property → composition → synthesis
  • 📚 Literature-grounded — retrieves and reasons over prior work
  • 🎲 Uncertainty-aware — ranks hypotheses by confidence
  • 🛑 Failure diagnosis — explains why experiments failed, proposes corrections
  • 🔄 Adaptive planning — updates strategy after each batch

Inner Loop (per cycle)

Step 1
Retrieve
Literature + past experiments
↓
Step 2
Reason
Chain-of-thought hypothesis generation
↓
Step 3
Rank & Dispatch
Select top-k, send to L2
↓
Step 4
Update Belief
Integrate results → next cycle
Design Philosophy

Purposeful Imperfection
applied to orchestration

In ML (our paper)

📊

Problem: Representational detail exceeds data fidelity → noise amplification

✅

Solution: Coarsen where unreliable, refine where stable

📈

Result: R² 0.04 → 0.88 on OER prediction

In NeverOT Orchestration

🤖

Problem: Primitive workflows assume deterministic agents → brittleness

✅

Solution: Match control granularity to agent reliability

🚀

Result: Robust closed-loop operation at scale

"The same epistemic principle that improves ML models also governs how we should design agent orchestrators: represent and control at the level of reliable knowledge."

Open Problems

Challenges Worth Solving Together

🎯

Goal Drift

In long-horizon tasks, SOTA agents inherit drift from weaker collaborators. arXiv:2603.03258 — only GPT-5.1 resists it.

🔍

Verifiable Reasoning

How do we audit why an agent proposed a hypothesis? Scientific reproducibility demands explainability.

⚡

Latency vs. Depth

Deeper reasoning per cycle vs. faster iteration rate — what's the right trade-off for discovery speed?

🔗

Multi-Lab Coordination

Can NeverOT-style agents coordinate across physically distributed labs, sharing knowledge without leaking IP?

Current status: Single-lab loop validated · 22 unit tests passing · L3 InverseDesignAgent v1 complete · Seeking collaborators for multi-lab extension

Thank You

Let's Build the
Never-Ending Lab Together

Key Takeaways

  • 🧱 Primitive workflows don't scale — agent-native does
  • 📐 Control granularity must match agent reliability
  • 🤝 Multi-agent hierarchy enables specialization + collaboration
  • 🔁 Never-ending optimization = science at machine speed

Get In Touch

Sissi Feng

Acceleration Consortium
University of Toronto

📄 Matter 9, 102652 (2026) — Purposeful Imperfection 🌐 sissifeng.github.io

Interested in collaborating?
Multi-lab coordination · Goal-drift robustness · Explainable hypothesis generation