My focus is on natural language processing, and increasingly on what lies beyond it: how large language models reason, how efficiently they do it, and whether that reasoning survives when the input is no longer text.

Before A*STAR, I did my Ph.D. at the University of Southern California with Prof. C.-C. Jay Kuo, and interned at Amazon working on LLMs for e-commerce with Karim Bouyarmane.

Research

My work spans the pipeline, from what a model is trained on, to how it decodes, to how we find out where it breaks.

LLM reasoning

Two questions here: one about ability, one about cost. CoinMath asks what coding data teaches mathematical reasoning. The style of code-based rationales, not just their correctness, shapes what a model learns, and diversifying those styles beats adding more general-domain code. InfoDensity asks how briefly a model can reason and still be right, rewarding information-dense traces so it reaches the answer in fewer tokens instead of padding its way there.

Decoding and structured generation

CABS estimates confidence at the sub-structure level rather than per token, then uses it to steer beam search and refine prompts. That cuts hallucination in structured data generation, where an entry can be perfectly well-formed and still wrong.

Evaluation and robustness

Spoken-MQA and IFEval-Audio ask whether mathematical reasoning and instruction-following survive in audio LLMs, rather than testing transcription alone. Resilience measures how far models degrade when instructions arrive with ASR slips, OCR errors, typos, or distracting content.

Selected work

InfoDensity

EMNLP 2026

Rewarding information-dense reasoning traces, so models reason efficiently instead of padding.

CoinMath

ACL 2025 Findings

How coding style in code-based rationales shapes math reasoning, and why diversifying that style beats adding more code.

Spoken-MQA

Benchmark · 2025

Can speech-based models do math? A benchmark for multi-faceted mathematical reasoning delivered as audio.

CABS

Amazon · 2024

Confidence-aware sub-structure beam search that cuts hallucination when LLMs generate structured data.

IFEval-Audio

AACL 2025

Can audio LLMs follow an instruction they hear? A benchmark built with my mentee Yiming Gao.

Resilience

EMNLP 2024 Findings

How far LLMs degrade when instructions arrive with ASR slips, OCR errors, typos, or distracting content.

Publications

Selected papers

The complete and most up-to-date list is on Google Scholar.

LLMs & Reasoning

InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning

Chengwei Wei, Jung-jae Kim, Longyin Zhang, Shengkai Chen, Nancy F. Chen

EMNLP 2026 paper

CoinMath: Harnessing the Power of Coding Instruction for Math LLMs

Chengwei Wei, Bin Wang, Jung-jae Kim, Guimei Liu, Nancy F. Chen

ACL 2025 Findings paper code + data + model

Confidence-Aware Sub-Structure Beam Search (CABS): Mitigating Hallucination in Structured Data Generation with Large Language Models

Chengwei Wei, Kee Kiat Koo, Amir Tavanaei, Karim Bouyarmane

Technical Report, 2024 paper demo

Resilience of Large Language Models for Noisy Instructions

Bin Wang, Chengwei Wei, Zhengyuan Liu, Geyu Lin, Nancy F. Chen

EMNLP 2024 Findings paper

CRAFT: Extracting and Tuning Cultural Instructions from the Wild

Bin Wang, Geyu Lin, Zhengyuan Liu, Chengwei Wei, Nancy F. Chen

C3NLP @ 2024 Best Paper Award paper code

Bias and Fairness in Chatbots: An Overview

Jintang Xue, Yun-Cheng Wang, Chengwei Wei, Xiaofeng Liu, Jonghye Woo, C.-C. Jay Kuo

APSIPA Trans. SIP, 2024 paper

An Overview on Generative AI at Scale with Edge-Cloud Computing

Yun-Cheng Wang, Jintang Xue, Chengwei Wei, C.-C. Jay Kuo

IEEE OJ-COMS, 2023 paper

An Overview of Language Models: Recent Developments and Outlook

Chengwei Wei, Yun-Cheng Wang, Bin Wang, C.-C. Jay Kuo

APSIPA Trans. SIP, 2023 paper

A Focused Study on Sequence Length for Dialogue Summarization

Bin Wang, Chen Zhang, Chengwei Wei, Haizhou Li

arXiv:2209.11910, 2022 paper code

Speech & Audio

Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models

Bin Wang, Xunlong Zou, Shuo Sun, Wenyu Zhang, Yingxu He, Zhuohan Liu, Chengwei Wei, Nancy F. Chen, AiTi Aw

ASRU 2025 paper code + data + model

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun, AiTi Aw

AACL 2025 paper code + data

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

Chengwei Wei, Bin Wang, Jung-jae Kim, Nancy F. Chen

arXiv:2505.15000, 2025 paper code + data

A Green Learning Approach to Spoofed Speech Detection

Chengwei Wei, Runqi Pang, C.-C. Jay Kuo

ICASSP 2024 paper

Fundamental NLP

Word Embedding Dimension Reduction via Weakly-Supervised Feature Selection

Jintang Xue, Yun-Cheng Wang, Chengwei Wei, C.-C. Jay Kuo, et al.

APSIPA Trans. SIP, 2024 paper

SynWMD: Syntax-aware Word Mover's Distance for Sentence Similarity Evaluation

Chengwei Wei, Bin Wang, C.-C. Jay Kuo

Pattern Recognition Letters, 2023 paper code

Task-specific Dependency-based Word Embedding Methods

Chengwei Wei, Bin Wang, C.-C. Jay Kuo

Pattern Recognition Letters, 2022 paper

Vision & Multimodal

Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia

Samuel Cahyawijaya, Holy Lovenia, …, Chengwei Wei, et al.

ACL 2025 paper

Efficient Human-Object-Interaction (EHOI) Detection via Interaction Label Coding and Conditional Decision

Tsung-Shan Yang, Yun-Cheng Wang, Chengwei Wei, Suya You, C.-C. Jay Kuo

CVIU, 2024 paper

GHOI: A Green Human-Object-Interaction Detector

Tsung-Shan Yang, Yun-Cheng Wang, Chengwei Wei, C.-C. Jay Kuo

IEEE MIPR 2024 paper

ExpressionHop: A Lightweight Human Facial Expression Classifier

Chengwei Wei, C.-C. Jay Kuo, Rafael Luiz Testa, Ariane Machado-Lima, Fátima L. S. Nunes

IEEE MIPR 2022 paper

Mentees

I have been fortunate to work with these students

Yiming Gao

Undergraduate Research Intern

NTU, Singapore

Jan 2025 – May 2025

Multimodal LLMs · co-advised with Bin Wang and Shuo Sun · AACL 2025

Pham The Binh Minh

Undergraduate Research Intern

NTU, Singapore

Jan 2025 – May 2025

Multimodal LLMs · co-advised with Bin Wang and Shuo Sun

Runqi Pang

Graduate Research Intern

University of Southern California

May 2023 – Aug 2023

NLP and speech processing · ICASSP 2024

Teaching

Teaching assistant & grader at USC

EE 559 Machine Learning I: Supervised Methods Spring 2023 · with Prof. Keith M. Chugg and Prof. B. Keith Jenkins
Math 229 Calculus III for Engineers and Scientists Fall 2021 and 2022 · with Prof. Guillaume Dreyer
Math 226 Calculus III Spring 2021 · with Prof. Cymra Haskell
EE 569 Introduction to Digital Image Processing (grader and mentor) Spring 2020 · with Prof. C.-C. Jay Kuo