AI researcher / founder

I am an AI researcher and founder working on autonomous AI research, generative AI, large language models, and synthetic data. I train LLMs and embedding models from scratch, largely using synthetic data, with models accumulating over 6M downloads on Hugging Face.

I hold a PhD in Machine Learning and Computer Science from the University of Tübingen (magna cum laude), and have 9+ years of experience across generative modeling, synthetic data, LLM alignment, and explainable AI. My research has received more than 3,000 citations, and my open-source work has been used by Google, AWS, the World Bank, and others.

Focus

My current work focuses on building AI systems that can autonomously explore, test, and improve solutions to hard research and optimization problems. I am particularly interested in systems where progress can be measured and independently verified.

  • AI agents and autonomous research
  • Generative AI and large language models
  • AI for algorithms and optimization
  • Synthetic data and representation learning
  • Efficient training and inference
  • LLM alignment: SFT, DPO, and GRPO
Selected work

Numaro

Autonomous AI research for hard problems in algorithms, optimization, mathematics, and scientific computing. Numaro generates, tests, and verifies candidate solutions while learning from previous experiments.

tuetoken

A high-performance tokenizer backend for LLM workloads, benchmarked up to 30× faster than tiktoken or Hugging Face tokenizers.

Faust-1

A 1.6B-parameter German language model trained from scratch, designed for efficient deployment on consumer hardware.

Synthetic Tabular Data

Research and open-source software for generating realistic synthetic tabular data with language models, used by Google (Kaggle), AWS, and practitioners across industry and research.

News

Published tuetoken

Released a high-performance tokenizer backend for LLM workloads.

Released Faust-1

Published a 1.6B-parameter German language model trained from scratch.

YapBench on arXiv

Introduced a benchmark for measuring verbosity and over-generation in conversational language models.

Contact

I am interested in research collaborations and ambitious problems in AI, algorithms, optimization, and scientific computing. For project or research discussions, feel free to get in touch.

vadim@tabularis.ai