Pooja Sethi

Hi, I'm Pooja.

I am an AI researcher building multimodal AI systems that understand and act in the physical world.

I recently joined Perceptron AI as a Member of Technical Staff, where I work on post-training for our vision-language model and harness engineering for embodied AI applications. Before that, I spent 8 years at Meta, most recently as a tech lead for the Machine Translation team.

I completed an M.S. in CS at Stanford (2021–2025), concentrating in Artificial Intelligence. Some sample projects include RL for instruction following (implemented DPO from scratch) and vision transformers for speech. Before that, I earned a B.S. in Computer Engineering at the University of Washington (2017).

The best way to reach me is via email.

Pooja Sethi outdoors beside a river

Research

I've always been drawn to getting models to better understand people, across languages and across modalities. I started with language, then moved to speech, and now work on video and understanding the physical world for robotics applications.

Writing

I write about AI research, and occasionally about life. All writing →

Personal

I strongly believe in building technology for human good. For health, ethical, and environmental reasons, I am vegan. I also like to read and occasionally write about things I'm learning. I also take photos, mostly of the Pacific Northwest. My name, पूजा, means prayer in Hindi.