Reasoning is expanding beyond text benchmarks into multimodal, spatial, and embodied settings. Models and agents must interpret long contexts, understand visual scenes, plan actions, use tools, and interact with changing environments, while meeting practical constraints on latency, memory, throughput, and serving cost. Yet stronger reasoning often depends on more test-time computation, longer trajectories, or deeper internal computation.
Meeting these challenges requires progress across the full pipeline, from model design to agent applications. Recurrent and latent architectures, reasoning-trace compression, reinforcement learning, adaptive model and budget selection, and KV-cache compression offer complementary routes to efficiency. In agent systems, these advances must work together with memory, tool use, and environment interaction: a cheaper model call is only useful if it also helps reduce the total cost of completing a task reliably.
The proposed third workshop brings together researchers and practitioners working on architectures, algorithms, training data, agent systems, evaluation, safety, and applications. Building on the NeurIPS 2025 and COLM 2026 editions, we aim to connect these perspectives and understand how capable reasoning models and agents can operate efficiently, robustly, and at scale under computational, memory, and interaction constraints.
When is additional computation worth its cost?
How can a model estimate the value of further reasoning, sampling, or verification? We seek policies that adapt to task difficulty, uncertainty, and available budgets, including the overhead of routing and auxiliary decision models.
What is the right representation for efficient reasoning?
Natural-language traces, recurrent or latent computation, and executable tools offer different trade-offs in cost, supervision, and control. Which representations support which reasoning operations, and when should a system switch between them?
What information must be preserved for reliable reasoning?
Compression can remove evidence, constraints, or intermediate conclusions needed later. How can reasoning-trace distillation, context compression, and agent memory retain decision-critical information and avoid delayed errors or repeated retrieval?
What training signals enable transferable reasoning efficiency?
Demonstrations, process feedback, failed attempts, and recovery trajectories shape how models allocate effort. Which signals teach reusable strategies that remain effective on unfamiliar tasks and under tighter computational budgets?
How can agents minimize the total cost of successful task completion?
An agent's cost includes reasoning, memory updates, tool calls, communication, and interaction with its environment. How can these decisions be optimized together while accounting for failures, retries, and expensive recovery?
How should reasoning efficiency be measured and compared?
Tokens, compute, latency, memory, energy, and monetary cost capture different trade-offs. Evaluations should consider full task trajectories, realistic workloads, variation across runs, tail latency, and the training or infrastructure costs behind inference-time savings.
How can efficiency improve while preserving robustness and safety?
Compression, quantization, distillation, and budget limits can affect uncertainty, error detection, and alignment. We ask when systems should reason further, seek evidence, abstain, or request human oversight, particularly under distribution shifts.