Learned surrogates enable rapid scenario screening by approximating the outputs of computationally expensive simulators, but decision support also demands traceable predictions, native routing confidence, and inspectable models. Soft mixtures of experts allow compact component pools, yet their predictions combine multiple experts. Even confident routing does not ensure that the dominant expert faithfully represents the deployed soft mixture. We introduce SDT-MoE, a soft-tree-gated mixture of experts that decouples routing resolution from the number of experts. A sparse soft tree provides fine-grained routing, while a data-driven leaf-to-expert map aggregates leaf probabilities into a distribution over shared linear response regimes. The dominant regime provides a hard per-prediction explanation, and its mass supplies a native routing-based confidence signal for selective prediction. A teacher--student construction clusters the leaf models of a full-granularity soft-tree teacher, initializes the shared experts from the resulting centroids, then fine-tunes the student with an expert-autonomy objective and sparsity-inducing hard-concrete split masks. An urban-flood emulation study in a real urban area compares surrogates from four representative families. SDT-MoE achieves competitive accuracy and the highest hard--soft fidelity among the evaluated soft-routed models, while providing inspectable routing and a compact global regime set.
Route Fine, Explain Compactly: A Soft-Tree-Gated Mixture-of-Experts Surrogate
Francesco Folino
;Luigi Pontieri
2026
Abstract
Learned surrogates enable rapid scenario screening by approximating the outputs of computationally expensive simulators, but decision support also demands traceable predictions, native routing confidence, and inspectable models. Soft mixtures of experts allow compact component pools, yet their predictions combine multiple experts. Even confident routing does not ensure that the dominant expert faithfully represents the deployed soft mixture. We introduce SDT-MoE, a soft-tree-gated mixture of experts that decouples routing resolution from the number of experts. A sparse soft tree provides fine-grained routing, while a data-driven leaf-to-expert map aggregates leaf probabilities into a distribution over shared linear response regimes. The dominant regime provides a hard per-prediction explanation, and its mass supplies a native routing-based confidence signal for selective prediction. A teacher--student construction clusters the leaf models of a full-granularity soft-tree teacher, initializes the shared experts from the resulting centroids, then fine-tunes the student with an expert-autonomy objective and sparsity-inducing hard-concrete split masks. An urban-flood emulation study in a real urban area compares surrogates from four representative families. SDT-MoE achieves competitive accuracy and the highest hard--soft fidelity among the evaluated soft-routed models, while providing inspectable routing and a compact global regime set.| File | Dimensione | Formato | |
|---|---|---|---|
|
ICTAI_2026.pdf
solo utenti autorizzati
Tipologia:
Documento in Pre-print
Licenza:
NON PUBBLICO - Accesso privato/ristretto
Dimensione
442.72 kB
Formato
Adobe PDF
|
442.72 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


