Anna Min | 闵安娜

Hi! I received B.E. at Tsinghua University, where I study computer science and statistics.

Email: anna.min1754@gmail.com (if you prefer an edu email, annamin@csail.mit.edu)

Portrait of Anna Min

Recent Updates

Research

I study intelligent agents through multimodal perception, generative models, and world modeling. My previous work explores how different modalities, including vision, audio, and language, interact and emerge in intelligent systems.

Selected Publications

Supervising Sound Localization by In-the-wild Egomotion

Supervising Sound Localization Using In-the-wild Ego-motion

Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens

CVPR 2025 (Highlight -- 2.98% accept rate) / [Paper]

Learn spatial sound sources in the wild using ego-motion signals derived from visual cues with limited perspectives.

A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation

A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation

Anna Min*, Chenxu Hu* , Yi Ren , Hang Zhao

Interspeech 2024 / [Paper]

Propose a dataset and pipeline with aligned bilingual audio tracks sharing similar emotions without using text as an intermediate for lesser-spoken languages and dialects.

Submissions

Physically-Grounded Video to Spatial Audio Generation

Kan Jen Cheng, Anna Min*, Tingle Li, Gopala Anumanchipalli

Submitting to ICLR 2027

Generate spatial audio from video by grounding sound generation in physical interactions and visual scene understanding.

Unified Spatial Audio Understanding And Generation

Kan Jen Cheng, Anna Min*, Tingle Li, Gopala Anumanchipalli

Submitting to ICLR 2027

A unified framework for spatial audio perception and generation, connecting multimodal understanding with controllable audio synthesis.

See, Infer, Intervene: Proactive World Modeling for Goal-Oriented Social Intelligence

Anna Min*, Honghui Zhang*, Chenmeinian Guo, Yichen Yu, Guanyu Liu, Yujia Zhang, Yongming Qin, Chongguo Song, Mengyue Yang, Lei Yu, Tianyu Shi

In submission to EMNLP 2026

Develop proactive world models that enable agents to infer social dynamics and intervene toward goal-oriented intelligence.

Projects

Pika 1.5

Anna Min (Core Contributor)

Generative Multi-modal Model

Core contributor to Pika 1.5, a generative video model enabling text-to-video creation and video editing.

Other Works

Quantifying Geometrical Associations Across Multi-modal Perception: An Information-theoretic Perspective

Anna Min, Hang Zhao

NeurIPS 2024 Wi@ML Workshop

Study geometric relationships across multimodal representations from an information-theoretic perspective.

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation

Anna Min, Chenxu Hu , Yi Ren , Hang Zhao

Preprint

Revisit cascaded speech translation systems and analyze when end-to-end approaches are unnecessary.

Selected Awards

Service

Misc: Art Creation

In the past I designed and developed video games (eg: A Player vs AI Strategy Game ). My Chinese name is pronounced as Ān Nà, which is similar to Anna.