Unified LLM Routing

Weaving Mode Selection and Execution into Unified LLM Routing

Three ways to route. One policy that chooses.

Xiaohan Wang1,2  Haozhen Zhang1  Qingyuan Liu1  Tao Feng3  Wenya Wang1

1Nanyang Technological University  ·  2Huazhong University of Science and Technology  ·  3University of Illinois Urbana-Champaign

single multi agentic one policy answer

The problem

No paradigm wins everywhere

Existing routers fix how model calls are organized before they ever see a query — and the best fixed choice is different in every domain.

single-round

One call, one answer

Cheapest and strongest among baselines on Reason and Recall — and badly outmatched on Code.

multi-round

Sequential calls

Each call reads the previous observation. Best among baselines on Math, Code and Knowledge; wasteful on a one-hop lookup.

agentic

Role-structured layers

Parallel calls with roles and references. Powerful where the work decomposes, overkill where it does not.

Which paradigm is strongest flips between domains, so any router that commits to one in advance has already given up accuracy somewhere. RouteWeaver makes the paradigm itself a routing decision.

Method

Make the paradigm a decision, then credit it separately

One Qwen3-4B policy writes the mode, then routes six frozen workers inside it — referring to them only by anonymous IDs with capability descriptions, never by model name or price.

RouteWeaver architecture: unified trajectory, COMET-GRPO and progressive exploration
One action grammar covers all three modes; COMET-GRPO splits credit between the mode decision and the execution inside it; progressive exploration hands mode selection over only once each mode can be executed.

One trajectory, three shapes

A rollout is (m, a₁, o₁, …, a_T, o_T, y): the mode, then worker calls interleaved with environment observations, then the answer. Single-round is exactly one call, multi-round two to four sequential calls, agentic up to eight calls over role- and reference-structured layers.

The mode, the actions and the answer are policy tokens. Worker output is context only — it never carries a gradient.

COMET-GRPO: two comparisons, one group

A low reward may mean the wrong mode, or the right mode executed badly. Ordinary GRPO cannot tell: it puts one trajectory-level advantage on every router token. COMET-GRPO splits it into an across-mode and a within-mode comparison, sharing one group σ so the two sum back to the original signal exactly.

A_mode = (μ_m − μ)/(σ+ε)    A_inner = (rᵢ − μ_m)/(σ+ε)

Learn to execute before deciding when

If the router chose modes from the start, an early preference would starve the others of experience. Progressive exploration forces the mode with probability ρ_b, annealed from all forced to all free, so the policy learns to execute each mode first and inherits the choice second.

ρ_b : 1 → 0   over   b₀ ≤ b < b₁

Results

Better than every fixed paradigm, in every domain

Nine benchmarks over five domains, against ten routers spanning all three execution paradigms.

61.4%average accuracy
+11.03%over the strongest baseline
5/5domains led
+18.28%on held-out benchmarks
MethodMathCodeKnowledgeReasonRecallAvg
Single-round routers
Router-KNN61.734.838.368.340.448.7
Router-MLP58.733.240.059.234.245.1
RouterDC61.329.434.257.535.043.5
GraphRouter64.636.134.269.239.648.7
Multi-round routers
Prompt LLM45.833.729.254.230.038.6
Router-KNN-MR40.833.230.848.329.636.6
R2-Reasoner45.830.728.350.023.335.6
Router-R166.352.848.968.739.855.3
Agentic routers
MasRouter54.237.641.762.538.346.8
GraphPlanner47.548.330.054.239.243.8
RouteWeaver 74.6+12.52% 58.2+10.23% 55.0+12.47% 76.7+10.84% 42.5+5.20% 61.4+11.03%

Accuracy (%), RouteWeaver trained with α = 0. Domain accuracy is the unweighted mean of its datasets; Avg is the unweighted mean of the five domains. Parentheses give the relative improvement over the strongest baseline in that column. Relative gains over the best single-round, multi-round and agentic baseline are 26.08%, 11.03% and 31.20%.

Both components matter — and one of them prevents collapse

VariantAvgS / M / A
RouteWeaver61.4.35 / .41 / .24
w/o COMET-GRPO58.8.13 / .82 / .05
w/o progressive exploration55.61.00 / .00 / .00

Without decision-specific credit the policy concentrates on multi-round. Without the forced-to-free schedule it collapses entirely to single-round and stops being a mode-selecting router at all. Both ablations still beat the strongest fixed-paradigm baseline.

It transfers to benchmarks held out of training

MethodAIMEGPQA-DBBHAvg
GraphRouter23.549.524.732.6
Router-R145.662.128.645.4
MasRouter28.143.828.233.4
RouteWeaver57.069.235.053.7

+11.4, +7.1 and +6.4 points over Router-R1. The adaptive behavior transfers too: agentic execution takes 75% of AIME queries but only 12% of BBH.

Routing behavior

It does not just pick the most expensive mode

The learned preferences differ sharply by domain — and in two of the five, almost every query goes to a single mode. Pick a domain:

single-round multi-round agentic

Share of test queries per execution mode (α = 0). Slices the paper does not label individually are shown unlabelled, so the bars still sum to 100%.

Accuracy and cost

One dial, four operating points

A second reward component scores worker cost, mixed in with weight α. Nothing else changes — same policy, same training, same pool.

α = 00.10.30.5
61.4% average accuracy

The reported model: task reward only, no efficiency term.

Cost is fixed-reference worker inference cost per query; at α = 0 it is reconstructed from token usage. Accuracy at α = 0 is 61.4%; the deltas are relative to it. The α > 0 routers are released as RouteWeaver-4B-cost.

Get it

Code, weights, data

The training and evaluation code is one repository with two entry points, on an unpatched verl 0.8. Evaluation and training share the same code path.

Reproduce the reported model

git clone https://github.com/LaughKing/RouteWeaver.git && cd RouteWeaver
bash setup_env.sh && conda activate routeweaver
python data/hf_download.py

bash scripts/train.sh cold     # phase 1: every query forced
bash scripts/train.sh main     # phases 2-3: forced to free, with COMET-GRPO

CKPT=... TAG=routeweaver GPU=0 OOD=1 bash scripts/eval.sh paper

Citation

@article{routeweaver2026,
  title   = {RouteWeaver: Weaving Mode Selection and Execution into Unified LLM Routing},
  author  = {Wang, Xiaohan and Zhang, Haozhen and Liu, Qingyuan and Feng, Tao and Wang, Wenya},
  journal = {arXiv preprint},
  year    = {2026}
}