Skip to content
IEEE white paper · HSI 2026 · Jul 17, 2026

Adaptive Inference Orchestration for Enterprise AI Assistants: Tiered Routing and Conditional Execution

18th IEEE International Conference on Human System Interaction
78.5% fewer unnecessary inference cycles (measured)
90.9% projected total cost reduction with tiered routing
Task completion accuracy above 95%
Abstract

The proliferation of Large Language Model (LLM)-powered personal assistants in enterprise environments introduces significant challenges in balancing response quality, latency, and inference cost, particularly in agentic workflows where multiple collaborating agents trigger cascading model calls. This paper presents an adaptive inference orchestration architecture for enterprise AI assistants that employs three key optimization strategies: (1) tiered model routing, which dynamically selects model complexity based on task classification; (2) trigger-based conditional execution, which uses lightweight deterministic gates to prevent unnecessary LLM invocations when no actionable context exists; and (3) hierarchical agent decomposition, which distributes workloads across specialized sub-agents with appropriate model tiers rather than routing all tasks through a single high-cost model. We evaluate the architecture in a real-world enterprise deployment over three weeks, integrating with Slack messaging, Microsoft Outlook email and calendar, and knowledge retrieval systems. Results demonstrate a 78.5% reduction in unnecessary inference cycles through conditional gating and a projected 90.9% total cost reduction when combined with tiered routing, while maintaining task completion accuracy above 95%. Our findings suggest that inference-aware orchestration design is essential for sustainable deployment of multi-agent AI systems in interactive enterprise contexts.

Index termsInference optimization · Multi-agent systems · Model routing · Enterprise AI · Agentic AI