Ph.D. student at ETH Zürich · Pre-Doctoral Researcher at IBM Research – Zurich
I work on multimodal machine learning for document understanding — My research focuses on multimodal machine learning for document understanding. I work within the Docling Team, focusing on Vision-Language Models for document parsing and chemical structure extraction.
At ETH I am part of the Computer Vision Lab led by Prof. Dr. Ender Konukoglu. At IBM Research I work with Dr. Peter W. J. Staar in the AI for Knowledge group, as part of the team behind Docling. I hold an M.Sc. in Robotics, Cognition, Intelligence from TU Munich.
Computer Vision · Multimodal Machine Learning · Vision-Language Models · Document AI · OCR · Optical Chemical Structure Recognition · AI for Chemistry
| Year | Work | Venue |
|---|---|---|
| 2026 | MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures — first author | CVPR 2026 |
| 2026 | Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images | ICDAR 2026 |
| 2025 | Advanced Layout Analysis Models for Docling | arXiv:2509.11720 |
Full list on Google Scholar.
I also make music WITHOUT AI :) My tracks are on Spotify:
📧 tstrohmeyer@ethz.ch (ETH) · Tim.Strohmeyer1@ibm.com (IBM) · tim.strohmeyer@t-online.de (personal)
Always happy to talk about document AI, chemistry-aware models, or collaborations.
