CAIS 2026
Meta Info
Homepage: https://www.caisconf.org
Paper list: https://www.caisconf.org/program/2026/papers/
Proceedings: https://dl.acm.org/doi/proceedings/10.1145/3786335
Papers
LLM Inference
XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs [Paper]
SJTU & CMU
Extend XGrammar with dynamic grammar support for efficient structured output generation in agentic LLM workloads (e.g., tool calling with runtime-defined schemas).
Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators [Paper]
Stanford
Replace KV cache with spectral Koopman operator estimation for constant-memory associative recall.
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints [Paper]
LinkedIn & MIT
Optimize LLM query routing at the batch level under joint cost and capacity constraints for multi-model serving.
Understanding and Improving Communication Performance in Multi-node LLM Inference [Paper]
UMD & LLNL
Characterize and optimize inter-node communication bottlenecks in multi-node LLM inference deployments.
Diffusion Model Inference
SwiftFusion: Scalable Sequence Parallelism for Distributed Inference of Diffusion Transformers on GPUs [Paper]
UofT & Amazon & NVIDIA & AWS
Introduce scalable sequence parallelism for distributed inference of Diffusion Transformers (DiTs) across multiple GPUs.
LLM Optimization
Scaling Textual Gradients via Sampling-Based Momentum [Paper]
UChicago & UT Austin & Santa Clara & Princeton & MSR & SylphAI
Scale textual gradient optimization (TextGrad) via sampling-based momentum for improved convergence.
Acronyms
DiT: Diffusion Transformer
KV: Key-Value
LLM: Large Language Model
LoRA: Low-Rank Adaptation
Last updated