Portrait of Mehrdad Fazli

Mehrdad Fazli

Ph.D. Candidate in Computer Science, George Mason University

I develop methods that make multimodal models more faithful to visual evidence, reducing hallucination through decoding, attention, and self-improvement.

I am a Ph.D. candidate at George Mason University, advised by Prof. Ziwei Zhu. My research focuses on multimodal AI, with an emphasis on hallucination in large vision-language models (LVLMs): characterizing why these models generate content that is not grounded in the visual input, and developing methods that improve their faithfulness. My broader work spans medical AI, explainable LLMs, and deep learning for health and sensing data.

My industry experience includes internships as a Ph.D. Software Engineering Intern at Google (Google Cloud), an Applied Scientist Intern at Amazon, and a Data Science and Machine Learning Intern at Wayfair. Prior to GMU, I received an M.Sc. in Systems Engineering from the University of Virginia.

Open to full-time roles as a Research Scientist, Applied Scientist, or Software Engineer, starting in 2027 (expected graduation: August 2027). I welcome inquiries at mfazli@gmu.edu.

News

Experience

Ph.D. Software Engineering Intern

Google, Google Cloud Platform, Persistent Disk
May 2026 – Aug 2026
Kirkland, WA
  • Designed and implemented a multi-agent AI framework to automate slow-IO debugging and root-cause analysis for GCP's backend storage, saving hundreds of engineering hours per year.
  • Triaged 20 production incidents over 3 weeks with the framework.

Applied Scientist Intern

Amazon, Personalization (P13N), Product Intent
Dec 2025 – Feb 2026
New York, NY
  • Fine-tuned LLaMA 3.2 3B and Qwen3 4B (LoRA) with Claude Sonnet 4.5 as a teacher to improve the faithfulness and personalization of explanations for recommended items.
  • The teacher–student distillation framework let student models match teacher performance at 3Γ— lower latency.

Data Science and Machine Learning Intern

Wayfair, Margin Optimization
Jun 2022 – Aug 2022
Boston, MA
  • Added confidence intervals to the demand simulator's short-term revenue and profit estimates to quantify uncertainty.
  • Built an evaluation framework on 50M+ transactions to retrospectively assess the simulator's accuracy (SQL, Python).
  • Created an error-decomposition framework to quantify the contribution of different error sources.

Researcher

University of Virginia
Aug 2019 – May 2023
Charlottesville, VA
  • Built deep probabilistic forecasting models (DeepAR, DeepTCN, TFT) showing that wastewater viral load improves COVID-19 forecasting.
  • Fit partially observed Markov process models to agent-based simulations calibrated to real COVID-19 prevalence data.
  • Modeled product diffusion over networks to find optimal prices under local network effects.
  • Teaching assistant for data and information engineering, differential equations, probability, and statistics.

Research

Multimodal AI and vision-language models

4 projects
Self-improving multimodal model overview
Current project

Self-Improving Multimodal Models for Improved Visual Understanding

Vision-language models often rely on language priors rather than attending closely to the image. I am building frameworks in which a model improves its own visual understanding by generating, verifying, and learning from visually grounded feedback, without relying on large amounts of new human annotation.

CAAC framework overview
WACV 2026

CAAC: Confidence-Aware Attention Calibration

A training-free method that corrects spatial perception bias and modality bias in LVLMs through visual-token calibration and confidence-guided attention re-scaling, reducing hallucination on CHAIR, AMBER, and POPE.

Context Embedding Injection overview
ACL 2026 Findings

Inject to Heal: Context Embedding Injection

Truthful tokens accumulate probability earlier across decoder layers than hallucinated ones. CEI uses the hidden state of the last input token to keep generation grounded in the image, with static and confidence-adaptive variants.

Faithfulness versus informativeness analysis
Under review

Does Playing it Safe Count as Faithfulness?

A diagnostic study of whether hallucination-mitigation methods genuinely improve grounding or simply decode more conservatively, evaluating six inference-time methods across three LVLMs and four benchmarks.

Medical and health AI

3 projects
VeriSim framework overview
EMNLP 2026 Findings

VeriSim: Stress-Testing Medical AI

A patient simulator that injects controllable, clinically grounded communication noise while verifying facts against UMLS, showing that diagnostic models lose 15–25% accuracy under realistic patient behavior.

COVID-19 forecasts from DeepTCN and TFT
IEEE BigData 2023

COVID-19 Forecasting with Deep Learning

Temporal convolutional networks and Temporal Fusion Transformers that use wastewater viral load and socio-economic data to improve COVID-19 forecasts.

Wastewater-based SEIR model
IEEE BigData 2021

Wastewater-Based COVID-19 Surveillance

Predicting community infections from wastewater RNA concentration with partially observed Markov processes and SEIR models, as part of the UVA COVID-19 surveillance program.

Explainable AI

1 project
ProtoSurE framework overview
🎀 AAAI 2026, Oral

ProtoSurE: Prototype-Based Explanations for LLMs

A prototype-based surrogate framework that provides faithful, sentence-level explanations for black-box LLM classifiers by matching inputs to human-understandable prototypes.

Human activity recognition and sensing

2 projects
HHAR-Net hierarchical classifier
πŸ† IHCI 2020, Best Paper

HHAR-Net: Hierarchical Activity Recognition

A two-level hierarchical neural classifier for smartphone and smartwatch sensor data that outperforms flat architectures.

NAPS Fusion overview
FUSION 2023

Sensor Fusion for Activity Recognition

A Dempster–Shafer-based fusion of smart-device sensors that builds adaptable, uncertainty-aware models of exercise and sedentary activity.

Operations research

1 project
Sales propagation model for two competing products
EJOR 2023

Price Optimization in a Network

A stochastic model of sales propagation among consumers for studying competitive and cooperative pricing under local network effects.

Publications

* Equal contribution

Education

Awards and honors

Service

Skills

Research areas
Multimodal LLMsVision-language modelsHallucination mitigationEvaluation and benchmarkingExplainable AI
Post-training
Supervised fine-tuningLoRA / PEFTKnowledge distillationReinforcement learning
Agentic AI
Multi-agent systemsTool useLLM-driven debugging and root-cause analysis
Frameworks
PyTorchHugging Face Transformersscikit-learn
Languages and cloud
PythonSQLC++BashGCP (Vertex AI, BigQuery)AWS (SageMaker)PySparkGit