There is something worth routing.
An oracle representation portfolio improves success while costing 13.7–35.3% less than the strongest fixed reference in the corresponding comparison.
The oracle says sometimes. The learned router says the harder part is predicting when. This project turns that gap into the result: a real routing opportunity can exist while the supervision needed to learn it is structurally scarce.
The project first proves that routing is worth attempting, then shows the learned policy fails under honest evaluation, and finally diagnoses where the learning problem breaks.
An oracle representation portfolio improves success while costing 13.7–35.3% less than the strongest fixed reference in the corresponding comparison.
Under true nested cross-validation, none of the six evaluated VWA cells Pareto-dominates the trivial always-cheapest policy.
A useful “which representation mattered?” label appears only after the task is solved. Low base success means useful supervision is scarce before classifier choice becomes the main issue.
The six-mode study factorizes text content, prompt style, and rendered image context so that multimodality is not treated as one monolithic switch.
When is expensive multimodal context truly necessary — and can that “when” be predicted cheaply enough not to erase the benefit?
Important: the phantom modes are experimental factorization tools, not claimed as universal cheap replacements for full SoM.
A failed router alone is weak evidence. The scientific story needs an oracle ceiling, honest controls, a failed learned policy, and a mechanism that explains the failure.
The oracle can choose representations task by task and beats fixed policies on success–cost trade-offs. Without this step, “the router failed” could simply mean there was nothing useful to route.
phase1_full_prereg_decision.jsonrouter_objective_ordering.mdPhantom arms uniquely solve subsets of tasks, but the project does not inflate that into a replacement claim. Same-condition rerun instability is large enough that small drop-one effects must be interpreted cautiously.
True nested cross-validation, a label-shuffle null, and the fixed always-cheapest baseline make the negative result meaningful. The learned policies can hit trade-off points, but none dominates the trivial policy.
router_triage_learnability.mdRouting labels are endogenous to task success: if a task is not solved, it cannot reveal which representation was necessary. At low base success, resplitting or tuning a classifier cannot manufacture missing solve events.
router_label_supply_diagnosis.mdThis is why the failure is not summarized as “the classifier was weak.” The scarcity sits upstream of classifier selection.
The training set is therefore filtered by the agent’s own competence. When solve events are rare, the label process becomes the bottleneck. More flexible classifiers do not fix missing events.
Live web benchmarks are stateful, authenticated, slow, and easy to contaminate. The infrastructure treats evaluation integrity as part of the research method, not backstage plumbing.
The goal is not simply to run many episodes. It is to make every scored claim traceable through resets, logs, scored-set accounting, contamination checks, and deterministic analysis entry points.
Negative results are easy to overstate in either direction. The portfolio keeps the claim ladder visible.
phantom_som is not a universal cheaper replacement for full SoM.Both are listed as submissions, not acceptances. The public page will be easy to update when decisions arrive.
Second Workshop for REsearch on Agent Language Models · EMNLP 2026.
Official workshop listing ↗Grounded and Faithful Vision-Language Models for Real-World Deployment · NeurIPS 2026.
Official workshop site ↗The demo is supporting material for the representation factorization. The scientific headline above is intentionally broader than the old “phantom mode wins” narrative.