2026

EvoMed-Agent: Self-Evolving Reasoners for Multimodal Clinical Sequences via Active Counterfactual Verification

Xing Han, Yuxin Wang, Cara Chen, Hsing-Huan Chung, Shijun Li, Gautham Gudur, Weichen Dai, Paul Pu Liang, Suchi Saria

In submission

A self-evolving, multi-agent clinical decision-support system in which a Solver, Proposer, and retrieval-based Verifier share one LLM backbone. It becomes process-aware by generating counterfactual what-if stress-tests from its own memory of past patients and improving on the resulting signal, with no external oracle model.

EvoMed-Agent: Self-Evolving Reasoners for Multimodal Clinical Sequences via Active Counterfactual Verification

Xing Han, Yuxin Wang, Cara Chen, Hsing-Huan Chung, Shijun Li, Gautham Gudur, Weichen Dai, Paul Pu Liang, Suchi Saria

In submission

A self-evolving, multi-agent clinical decision-support system in which a Solver, Proposer, and retrieval-based Verifier share one LLM backbone. It becomes process-aware by generating counterfactual what-if stress-tests from its own memory of past patients and improving on the resulting signal, with no external oracle model.

On the Invariance and Generality of Neural Scaling Laws
On the Invariance and Generality of Neural Scaling Laws

Xing Han, Ziyin Liu, Suchi Saria, Paul Pu Liang

Preprint Top ~3% of NeurIPS submissions by reviewer rating

Shows that neural scaling laws are preserved under bijective, information-preserving data transformations and shift predictably otherwise, giving a principled rule for how model capacity should be chosen as training data varies in size and quality.

On the Invariance and Generality of Neural Scaling Laws

Xing Han, Ziyin Liu, Suchi Saria, Paul Pu Liang

Preprint Top ~3% of NeurIPS submissions by reviewer rating

Shows that neural scaling laws are preserved under bijective, information-preserving data transformations and shift predictably otherwise, giving a principled rule for how model capacity should be chosen as training data varies in size and quality.

MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling

Hsing-Huan Chung, Shijun Li, Yoav Wald, Xing Han, Suchi Saria, Joydeep Ghosh

Preprint

Encodes irregular multimodal clinical series as time-ordered XML triplets and fine-tunes an LLM in two stages, starting from value-redacted data, so the model learns from when and what clinicians choose to measure rather than recorded values alone.

MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling

Hsing-Huan Chung, Shijun Li, Yoav Wald, Xing Han, Suchi Saria, Joydeep Ghosh

Preprint

Encodes irregular multimodal clinical series as time-ordered XML triplets and fine-tunes an LLM in two stages, starting from value-redacted data, so the model learns from when and what clinicians choose to measure rather than recorded values alone.

FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning
FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

Xing Han, Shravan Chaudhari, Tanvi Ranade, Rama Chellappa, Suchi Saria

Preprint

A fixed-capacity mixture-of-experts framework with modality-specific routers that pretrains across flexible modality combinations and continually absorbs new tasks by compressing accumulated expert knowledge into low-rank memory subspaces, alleviating catastrophic forgetting.

FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

Xing Han, Shravan Chaudhari, Tanvi Ranade, Rama Chellappa, Suchi Saria

Preprint

A fixed-capacity mixture-of-experts framework with modality-specific routers that pretrains across flexible modality combinations and continually absorbs new tasks by compressing accumulated expert knowledge into low-rank memory subspaces, alleviating catastrophic forgetting.

Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts
Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts

Xing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang, Suchi Saria

International Conference on Learning Representations (ICLR) 2026 Top ~10% of accepted papers by reviewer rating

Quantifies pairwise temporal delays across many modalities and routes tokens by interaction type -- redundancy, uniqueness, and synergy -- so experts specialize in interaction patterns that remain interpretable and align with known physiology.

Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts

Xing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang, Suchi Saria

International Conference on Learning Representations (ICLR) 2026 Top ~10% of accepted papers by reviewer rating

Quantifies pairwise temporal delays across many modalities and routes tokens by interaction type -- redundancy, uniqueness, and synergy -- so experts specialize in interaction patterns that remain interpretable and align with known physiology.

QoQ-Med3: Robust Multimodal Clinical Analysis Foundation Model with Reasoning
QoQ-Med3: Robust Multimodal Clinical Analysis Foundation Model with Reasoning

David Dai, Jeannie She, Jiaee Cheong, Xing Han, Carl Harris, Haowen Wei, Farzan Vahedifard, Suchi Saria, Robert Stevens, Paul Pu Liang

npj Digital Medicine

A robust multimodal clinical foundation model with reasoning that integrates a wide range of clinical modalities and amplifies underrepresented ones such as ultrasound and mammography, yielding transferable performance across modalities, tasks, and institutions.

QoQ-Med3: Robust Multimodal Clinical Analysis Foundation Model with Reasoning

David Dai, Jeannie She, Jiaee Cheong, Xing Han, Carl Harris, Haowen Wei, Farzan Vahedifard, Suchi Saria, Robert Stevens, Paul Pu Liang

npj Digital Medicine

A robust multimodal clinical foundation model with reasoning that integrates a wide range of clinical modalities and amplifies underrepresented ones such as ultrasound and mammography, yielding transferable performance across modalities, tasks, and institutions.

2025

WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales
WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales

Drew Prinster#, Xing Han#, Anqi Liu, Suchi Saria (# corresponding author)

International Conference on Machine Learning (ICML) 2025

Introduces weighted conformal test martingales for anytime-valid post-deployment monitoring, adapting online to benign covariate shift while flagging harmful shifts and diagnosing whether the cause is covariate, concept, or out-of-support.

WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales

Drew Prinster#, Xing Han#, Anqi Liu, Suchi Saria (# corresponding author)

International Conference on Machine Learning (ICML) 2025

Introduces weighted conformal test martingales for anytime-valid post-deployment monitoring, adapting online to benign covariate shift while flagging harmful shifts and diagnosing whether the cause is covariate, concept, or out-of-support.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

Xiaojun Shan*, Qi Cao*, Xing Han*, Haofei Yu, Paul Pu Liang (* equal contribution)

Preprint

Groups instruction-tuning tasks by the type of multimodal interaction they demand -- redundancy, unique-modality dominance, or synergistic fusion -- avoiding the interference that makes naive scaling of task count fail.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

Xiaojun Shan*, Qi Cao*, Xing Han*, Haofei Yu, Paul Pu Liang (* equal contribution)

Preprint

Groups instruction-tuning tasks by the type of multimodal interaction they demand -- redundancy, unique-modality dominance, or synergistic fusion -- avoiding the interference that makes naive scaling of task count fail.

Between Linear and Sinusoidal: Rethinking the Time Encoder in Dynamic Graph Learning

Hsing-Huan Chung, Shravan Chaudhari, Xing Han, Yoav Wald, Suchi Saria, Joydeep Ghosh

Transactions on Machine Learning Research (TMLR)

Shows that the standard sinusoidal time encoder in dynamic graph learning discards information, and that a simple linear encoder lets self-attention learn time spans itself, with consistent gains across benchmarks.

Between Linear and Sinusoidal: Rethinking the Time Encoder in Dynamic Graph Learning

Hsing-Huan Chung, Shravan Chaudhari, Xing Han, Yoav Wald, Suchi Saria, Joydeep Ghosh

Transactions on Machine Learning Research (TMLR)

Shows that the standard sinusoidal time encoder in dynamic graph learning discards information, and that a simple linear encoder lets self-attention learn time spans itself, with consistent gains across benchmarks.

SSL-DA: Semi- and Self-Supervised Learning with Dual Attention for Echocardiogram Segmentation

Lin Lv, Xing Han, Zhengxiang Sun, Zhaoguang Li, Xiuying Wang, Tong Jiang, Yiren Liu, Tianshu Li, Jingjing Xu, Liangzhen You, Guihua Yao, Feng-rong Sun, Jianping Xing

Journal of Imaging Informatics in Medicine

A two-phase framework pairing self-supervised pre-training with temporal masking and semi-supervised segmentation with dual attention, enabling accurate left-ventricle segmentation in echocardiogram video from only sparse annotations.

SSL-DA: Semi- and Self-Supervised Learning with Dual Attention for Echocardiogram Segmentation

Lin Lv, Xing Han, Zhengxiang Sun, Zhaoguang Li, Xiuying Wang, Tong Jiang, Yiren Liu, Tianshu Li, Jingjing Xu, Liangzhen You, Guihua Yao, Feng-rong Sun, Jianping Xing

Journal of Imaging Informatics in Medicine

A two-phase framework pairing self-supervised pre-training with temporal masking and semi-supervised segmentation with dual attention, enabling accurate left-ventricle segmentation in echocardiogram video from only sparse annotations.

2024

FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion

Xing Han, Huy Nguyen, Carl Harris, Nhat Ho, Suchi Saria

Neural Information Processing Systems (NeurIPS) 2024 Over 150 citations as of July 2026

A mixture-of-experts Transformer whose Laplace gating fuses an arbitrary and variable number of asynchronous modalities while handling missing data and irregular sampling, backed by a convergence-rate guarantee.

FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion

Xing Han, Huy Nguyen, Carl Harris, Nhat Ho, Suchi Saria

Neural Information Processing Systems (NeurIPS) 2024 Over 150 citations as of July 2026

A mixture-of-experts Transformer whose Laplace gating fuses an arbitrary and variable number of asynchronous modalities while handling missing data and irregular sampling, backed by a convergence-rate guarantee.

Quadratic Gating Functions in Mixture of Experts: A Statistical Insight

Pedram Akbarian*, Huy Nguyen*, Xing Han*, Nhat Ho (* equal contribution)

In submission to SIAM Journal on Mathematics of Data Science

Proves that each row of self-attention can be written as a quadratic-gating mixture of linear experts, and derives convergence rates for expert estimation under quadratic gating.

Quadratic Gating Functions in Mixture of Experts: A Statistical Insight

Pedram Akbarian*, Huy Nguyen*, Xing Han*, Nhat Ho (* equal contribution)

In submission to SIAM Journal on Mathematics of Data Science

Proves that each row of self-attention can be written as a quadratic-gating mixture of linear experts, and derives convergence rates for expert estimation under quadratic gating.

On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions

Huy Nguyen*, Xing Han*, Carl Harris, Nhat Ho, Suchi Saria (* equal contribution)

In submission to JMLR

Replaces softmax with a Laplace gating function at both levels of a hierarchical mixture of experts, provably eliminating the harmful parameter interactions that slow expert estimation.

On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions

Huy Nguyen*, Xing Han*, Carl Harris, Nhat Ho, Suchi Saria (* equal contribution)

In submission to JMLR

Replaces softmax with a Laplace gating function at both levels of a hierarchical mixture of experts, provably eliminating the harmful parameter interactions that slow expert estimation.

Novel Node Category Detection Under Subpopulation Shift

Hsing-Huan Chung, Shravan Chaudhari, Yoav Wald, Xing Han, Joydeep Ghosh

European Conference on Machine Learning and Data Mining (ECML-PKDD) 2024

RECO-SLIP detects previously unseen node categories under subpopulation shift by combining recall-constrained optimization with selective link prediction.

Novel Node Category Detection Under Subpopulation Shift

Hsing-Huan Chung, Shravan Chaudhari, Yoav Wald, Xing Han, Joydeep Ghosh

European Conference on Machine Learning and Data Mining (ECML-PKDD) 2024

RECO-SLIP detects previously unseen node categories under subpopulation shift by combining recall-constrained optimization with selective link prediction.

Achieving Fairness Across Local and Global Models in Federated Learning

Disha Makhija, Xing Han, Joydeep Ghosh, Yejin Kim

Preprint

EquiFL adds a fairness term to each client's local objective together with a coordination mechanism that blocks bias from propagating during aggregation, improving fairness at both the local and global level.

Achieving Fairness Across Local and Global Models in Federated Learning

Disha Makhija, Xing Han, Joydeep Ghosh, Yejin Kim

Preprint

EquiFL adds a fairness term to each client's local objective together with a coordination mechanism that blocks bias from propagating during aggregation, improving fairness at both the local and global level.

2023

Designing Robust Transformers using Robust Kernel Density Estimation
Designing Robust Transformers using Robust Kernel Density Estimation

Xing Han, Tongzheng Ren, Tan Minh Nguyen, Khai Nguyen, Joydeep Ghosh, Nhat Ho

37th Conference on Neural Information Processing Systems (NeurIPS 2023)

Reformulates self-attention using robust kernel density estimation, yielding a family of attention mechanisms that down-weight contaminated samples and plug into diverse Transformer architectures.

Designing Robust Transformers using Robust Kernel Density Estimation

Xing Han, Tongzheng Ren, Tan Minh Nguyen, Khai Nguyen, Joydeep Ghosh, Nhat Ho

37th Conference on Neural Information Processing Systems (NeurIPS 2023)

Reformulates self-attention using robust kernel density estimation, yielding a family of attention mechanisms that down-weight contaminated samples and plug into diverse Transformer architectures.

Efficient Forecasting of Large Scale Hierarchical Time Series via Multilevel Clustering

Xing Han, Tongzheng Ren, Jing Hu, Joydeep Ghosh, Nhat Ho

9th International conference on Time Series and Forecasting

A multilevel clustering approach that combines Wasserstein distance with Soft-DTW divergence to cluster series jointly at local and global levels, then forecasts bottom-up for large hierarchies.

Efficient Forecasting of Large Scale Hierarchical Time Series via Multilevel Clustering

Xing Han, Tongzheng Ren, Jing Hu, Joydeep Ghosh, Nhat Ho

9th International conference on Time Series and Forecasting

A multilevel clustering approach that combines Wasserstein distance with Soft-DTW divergence to cluster series jointly at local and global levels, then forecasts bottom-up for large hierarchies.

A Novel Control-Variates Approach for Performative Gradient-Based Learners with Missing Data

Xing Han, Jing Hu, Joydeep Ghosh

2023 International Joint Conference on Neural Networks (IJCNN)

A control-variates correction that turns any imputation model, even a biased one, into unbiased gradient estimates for models learned on data with missing values, with proven improvements in SGD convergence.

A Novel Control-Variates Approach for Performative Gradient-Based Learners with Missing Data

Xing Han, Jing Hu, Joydeep Ghosh

2023 International Joint Conference on Neural Networks (IJCNN)

A control-variates correction that turns any imputation model, even a biased one, into unbiased gradient estimates for models learned on data with missing values, with proven improvements in SGD convergence.

2022

Dynamic Combination of Heterogeneous Models for Hierarchical Time Series

Xing Han, Jing Hu, Joydeep Ghosh

ICDM 2022 Workshop

DYCHEM dynamically combines heterogeneous expert forecasters per series, learns the aggregation hierarchy during training, and produces coherent probabilistic forecasts across it.

Dynamic Combination of Heterogeneous Models for Hierarchical Time Series

Xing Han, Jing Hu, Joydeep Ghosh

ICDM 2022 Workshop

DYCHEM dynamically combines heterogeneous expert forecasters per series, learns the aggregation hierarchy during training, and produces coherent probabilistic forecasts across it.

Machine-learning based generation of text style variations for digital content items

Jessica Lundin, Owen Winne Schoppe, Xing Han, Michael Reynolds Sollami, Brian J. Lonsdorf, Alan Martin Ross, David J. Woodward, Sonke Rohde

US Patent 2022/0245322 A1

An online system that generates style variations of a reference content item by applying machine-learned style transfer models, so digital content can be re-rendered in different textual styles.

Machine-learning based generation of text style variations for digital content items

Jessica Lundin, Owen Winne Schoppe, Xing Han, Michael Reynolds Sollami, Brian J. Lonsdorf, Alan Martin Ross, David J. Woodward, Sonke Rohde

US Patent 2022/0245322 A1

An online system that generates style variations of a reference content item by applying machine-learned style transfer models, so digital content can be re-rendered in different textual styles.

Architecture Agnostic Federated Learning for Neural Networks
Architecture Agnostic Federated Learning for Neural Networks

Disha Makhija, Xing Han, Nhat Ho, Joydeep Ghosh

Proceedings of the 39th International Conference on Machine Learning (ICML) 2022

FedHeNN lets federated clients train personalized models of any architecture, coordinating them through instance-level representations shared across peers rather than a common model or gradients.

Architecture Agnostic Federated Learning for Neural Networks

Disha Makhija, Xing Han, Nhat Ho, Joydeep Ghosh

Proceedings of the 39th International Conference on Machine Learning (ICML) 2022

FedHeNN lets federated clients train personalized models of any architecture, coordinating them through instance-level representations shared across peers rather than a common model or gradients.

2021

Multi-Pair Text Style Transfer for Unbalanced Data via Task-Adaptive Meta-Learning

Xing Han, Jessica Lundin

Proceedings of the 1st ACL Workshop on Meta Learning and Its Applications to Natural Language Processing

A task-adaptive meta-learning framework that performs multi-pair text style transfer with a single model, adaptively balancing meta-knowledge across highly unbalanced style pairs.

Multi-Pair Text Style Transfer for Unbalanced Data via Task-Adaptive Meta-Learning

Xing Han, Jessica Lundin

Proceedings of the 1st ACL Workshop on Meta Learning and Its Applications to Natural Language Processing

A task-adaptive meta-learning framework that performs multi-pair text style transfer with a single model, adaptively balancing meta-knowledge across highly unbalanced style pairs.

Model-Agnostic Explanations using Minimal Forcing Subsets

Xing Han, Joydeep Ghosh

2021 International Joint Conference on Neural Networks (IJCNN)

An iterative constrained-optimization algorithm that finds the minimal forcing subset of training samples whose removal would flip a model's decision, giving compact model-agnostic explanations.

Model-Agnostic Explanations using Minimal Forcing Subsets

Xing Han, Joydeep Ghosh

2021 International Joint Conference on Neural Networks (IJCNN)

An iterative constrained-optimization algorithm that finds the minimal forcing subset of training samples whose removal would flip a model's decision, giving compact model-agnostic explanations.

Split Localized Conformal Prediction

Xing Han, Ziyang Tang, Joydeep Ghosh, Qiang Liu

ICML 2021 DFUQ Workshop

A modified non-conformity score that uses kernel density estimation to locally approximate the conditional distribution, tightening prediction intervals while retaining the simplicity and coverage guarantee of split conformal prediction.

Split Localized Conformal Prediction

Xing Han, Ziyang Tang, Joydeep Ghosh, Qiang Liu

ICML 2021 DFUQ Workshop

A modified non-conformity score that uses kernel density estimation to locally approximate the conditional distribution, tightening prediction intervals while retaining the simplicity and coverage guarantee of split conformal prediction.

Simultaneously Reconciled Quantile Forecasting of Hierarchically Related Time Series
Simultaneously Reconciled Quantile Forecasting of Hierarchically Related Time Series

Xing Han, Sambarta Dasgupta, Joydeep Ghosh

Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS) 2021

A nonlinear model trained with quantile regression loss and coherency regularization that produces probabilistic forecasts for hierarchically related time series while keeping them consistent across aggregation levels.

Simultaneously Reconciled Quantile Forecasting of Hierarchically Related Time Series

Xing Han, Sambarta Dasgupta, Joydeep Ghosh

Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS) 2021

A nonlinear model trained with quantile regression loss and coherency regularization that produces probabilistic forecasts for hierarchically related time series while keeping them consistent across aggregation levels.

2020

Certified Monotonic Neural Networks
Certified Monotonic Neural Networks

Xingchao Liu, Xing Han, Na Zhang, Qiang Liu

34th Conference on Neural Information Processing Systems (NeurIPS 2020) Spotlight presentation (280/9454 ~ 2.96%); over 185 citations as of July 2026

Certifies monotonicity of general piecewise-linear networks via mixed-integer linear programming, so monotonicity can be encouraged heuristically during training and then verified exactly afterwards.

Certified Monotonic Neural Networks

Xingchao Liu, Xing Han, Na Zhang, Qiang Liu

34th Conference on Neural Information Processing Systems (NeurIPS 2020) Spotlight presentation (280/9454 ~ 2.96%); over 185 citations as of July 2026

Certifies monotonicity of general piecewise-linear networks via mixed-integer linear programming, so monotonicity can be encouraged heuristically during training and then verified exactly afterwards.

Transparent Interpretation with Knockout

Xing Han, Yihao Feng, Na Zhang, Qiang Liu

ICML Workshop on Human Interpretability in Machine Learning (WHI) 2020 Spotlight

Transparent Interpretation with Knockout

Xing Han, Yihao Feng, Na Zhang, Qiang Liu

ICML Workshop on Human Interpretability in Machine Learning (WHI) 2020 Spotlight

2019

Sensing Personality to Predict Job Performance

Suwen Lin, Stephen M. Mattingly, Xing Han

CHI Future of Work Workshop 2019

Applies machine learning to continuously collected wearable and smartphone sensor data to infer personality traits, and uses them to help predict individual job performance in the workplace.

Sensing Personality to Predict Job Performance

Suwen Lin, Stephen M. Mattingly, Xing Han

CHI Future of Work Workshop 2019

Applies machine learning to continuously collected wearable and smartphone sensor data to infer personality traits, and uses them to help predict individual job performance in the workplace.

2018

Effects of Integrated Intent Recognition and Communication on Human-Robot Collaboration

A. Gutierrez, M. L. Chang, Xing Han, K. C. Chang

International Conference on Intelligent Robots and Systems (IROS) 2018

Couples a robot's recognition of its human partner's motion intent with legible, predictable motion that signals the robot's own intent, and shows in a within-subjects user study that closing this bi-directional loop produces more collaborative team behavior.

Effects of Integrated Intent Recognition and Communication on Human-Robot Collaboration

A. Gutierrez, M. L. Chang, Xing Han, K. C. Chang

International Conference on Intelligent Robots and Systems (IROS) 2018

Couples a robot's recognition of its human partner's motion intent with legible, predictable motion that signals the robot's own intent, and shows in a within-subjects user study that closing this bi-directional loop produces more collaborative team behavior.