Hi, I'm Pooja.
I am an AI researcher building multimodal AI systems that understand and act in the physical world.
I recently joined Perceptron AI as a Member of Technical Staff, where I work on post-training for our vision-language model and harness engineering for embodied AI applications. Before that, I spent 8 years at Meta, most recently as a tech lead for the Machine Translation team.
I completed an M.S. in CS at Stanford (2021–2025), concentrating in Artificial Intelligence. Some sample projects include RL for instruction following (implemented DPO from scratch) and vision transformers for speech. Before that, I earned a B.S. in Computer Engineering at the University of Washington (2017).
The best way to reach me is via email.
Research
I've always been drawn to getting models to better understand people, across languages and across modalities. I started with language, then moved to speech, and now work on video and understanding the physical world for robotics applications.
- 2026–Mk1.5 · Perceptron AI
Egocentric video understanding and tool use for our flagship vision-language model for physical AI.
- 2022–26Multimodal Speech Translation · Meta AI (PyTorch | AI & Data Infra)
Tech lead for voice dubbing across languages, announced by Mark Zuckerberg at Meta Connect.
- 2021–22Document Understanding · Impira
Deep learning for document understanding, combining text and visual features.
- 2017–21
- 2014–17Respeak · UW
Crowdsourced speech transcription app deployed in India. Winner of UW's Best Undergraduate Honors Thesis Award.
Writing
I write about AI research, and occasionally about life. All writing →
Personal
I strongly believe in building technology for human good. For health, ethical, and environmental reasons, I am vegan. I also like to read and occasionally write about things I'm learning. I also take photos, mostly of the Pacific Northwest. My name, पूजा, means prayer in Hindi.