RUBICON: Do Language Models Know When They've Crossed the Point of No Return?
Testing whether long-horizon agents recognize when a task has become impossible, and report it to the user.
Hi! I'm Adnan. I currently work at IBM Research as a Research Engineer in the Data for AI team. I work on self-improving retrieval systems and explore how the systems around a model can help it work better.
Beyond my core role, I'm interested in building long-horizon agents, with a particular focus on making good data and creating meaningful metrics to measure their progress. I also enjoy probing frontier models to understand their limits and identifying the specific gaps where their reasoning or reliability starts to drop off.
In the past, I worked on designing evals for multimodal understanding and robustness. I also spent some time researching LLM jailbreaking and security.
I'm always open to discussing research, new opportunities, or just exchanging ideas. Feel free to reach out!
Testing whether long-horizon agents recognize when a task has become impossible, and report it to the user.
Efficient offline exploration of multi-index retrieval methods.
Consistency and robustness in chart understanding.
Structured insights from reviews and seller descriptions.
Reasoning across related charts.
LLMs plus static rules for code-data profiling.