Nvidia Corporation
Santa Clara, CA
NVIDIA is at the forefront of the AI revolution, and our research is shaping the future of large language models. We are looking for a Principal Scientist to set the technical direction for synthetic data generation across NVIDIA's frontier model efforts. You will define and build open-source libraries within the NVIDIA NeMo ecosystem that generate synthetic datasets across text, code, structured, and multimodal data, feeding the pre- and post-training of LLMs such as Nemotron. This role combines hands-on software engineering with applied research in generative methods, and you will collaborate with research, engineering, product, and model teams as well as external labs. What you'll be doing: Build and scale data generation pipelines using LLM-based methods combined with automated quality evaluation. resulting in datasets to improve both initial training and fine-tuning of LLMs such as Nemotron. These data pipelines cover reasoning, coding, structured output, and multimodal...