Syed Mumtahin Mahmud (Dipto)

I am a Computer Science and Engineering graduate from the University of Dhaka. My research centers on the reasoning capabilities of Large Language Models (LLMs), with a particular focus on structured, symbolic, and computational reasoning. My recent work investigates whether LLMs can synthesize formal computational artifacts, such as automata, from natural language and formal specifications, using this as a controlled and verifiable testbed for reasoning ability beyond surface-level language generation. This direction builds on a strong foundation in deep learning developed through prior research in computer vision and image restoration.
Currently, I am a Research Assistant at the Cognitive Agents and Interaction Lab at the University of Dhaka, advised by Dr. Md Mosaddek Khan, where I work on projects related to LLM reasoning and evaluation, Graph Neural Networks, and medical imaging.

Email  /  Scholar  /  LinkedIn  /  Github

profile photo

Research

My research interests lie in natural language processing and large language models, with a focus on structured computational reasoning, formal reasoning benchmarks, and LLM evaluation. A central line of my work asks whether LLMs can move beyond describing computation to actually constructing it, introducing the task of synthesizing pushdown automata from context-free language specifications and showing that even state-of-the-art models struggle to construct executable automata reliably, with performance degrading sharply as language complexity increases. I am broadly interested in extending this line of inquiry to other formal and symbolic reasoning tasks, as well as to the evaluation of LLM reasoning in educational and human-centered contexts.

My prior research experience spans computer vision and deep learning, including large-scale dataset construction for image deblurring, Vision Transformer and frequency-domain architectures for image restoration, medical imaging, and Graph Neural Networks, as well as educational data mining through large-scale analysis of iterative programming submission traces. This background continues to inform how I approach rigorous, evaluation-driven research design.

Education

University of Dhaka Logo University of Dhaka

BS.c in Computer Science and Engineering, (CGPA 3.77/4.00), Position: 5th, top 10% of class,

Jan 2020 - March 2025

Publications

PDA Synthesis
Can LLMs Design Computational Machines? Pushdown Automaton Synthesis as a Test of Structured Computational Reasoning
Syed Mumtahin Mahmud, Nazira Jesmin Lina, Shahriyar Zaman Ridoy, S. M. Muhtasimul Hasan, Md Mosaddek Khan
Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026

We ask whether LLMs can build formal computational machines, not just describe them. We introduce the task of pushdown automaton synthesis from context-free language specifications, with a benchmark of 41 CFLs and 50 labeled test strings each, and an automated framework that separates structural buildability from execution-based correctness. Evaluating six state-of-the-art models, we find that even the strongest system reaches only 75.6% accuracy while others fall below 15%, with logical errors dominating and failures rising sharply with language complexity. PDA synthesis thus offers a controlled, verifiable testbed for structured symbolic reasoning.

Code Verdict Prediction
Can LLMs Judge Code Without Execution? Predicting Code-Submission Verdicts from Problem Descriptions and Source Code
Syed Mumtahin Mahmud, Nazira Jesmin Lina, Nafiul Hasan Anik, Md. Fahim Arefin, Md Mosaddek Khan
Under Review, AAAI Conference on Artificial Intelligence (AAAI), 2027

We investigate whether LLMs can judge code correctness without executing it, formulating execution-free online-judge verdict prediction as a five-class task given only a problem description and submitted code. Using a benchmark of 5,729 human-written submissions spanning 46 problems and three programming languages, we evaluate seven frontier LLMs under zero-shot, few-shot, and chain-of-thought prompting, and further fine-tune smaller open-weight models with QLoRA under a strict problem-level split. Chain-of-thought prompting yields the strongest results, with the best model reaching 76.87% accuracy, while models recognize verdicts supported by visible source-level evidence far more reliably than execution-dependent failures such as runtime errors and time-limit violations.

Sharp Image
SloMoDeblur: A Large-Scale Smartphone Image Deblurring Dataset
Syed Mumtahin Mahmud, Mahdi Mohd Hossain Noki, Prothito Shovon Majumder, Abdul Mohaimen Al Radi, Sudipto Das Sukanto, Afia Lubaina, Md Mosaddek Khan
Journal of Data-centric Machine Learning Research (DMLR), 2026

We present the largest real-world image deblurring dataset, built from smartphone slow-motion videos. By averaging 240 fps frames to create blur and using the center frame as the sharp reference, we generate over 42,000 high-resolution blur-sharp pairs. This makes it roughly 10 times larger and 8 times more diverse than existing datasets. Covering a wide range of indoor and outdoor scenes with various object and camera motions, our benchmark reveals significant performance drops in state-of-the-art models, highlighting its complexity. The dataset and generation scripts are available on HuggingFace.

Sharp Image
From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring
Syed Mumtahin Mahmud, Mahdi Mohd Hossain Noki, Prothito Shovon Majumder, Abdul Mohaimen Al Radi, Md. Haider Ali Md Mosaddek Khan,
Proceedings of the 18th International Conference on Agents and Artificial Intelligence (ICAART), 2026

We propose a novel approach to blind image deblurring that combines Vision Transformers with Fast Fourier Transform (FFT) and ReLU sparsity priors. Our method efficiently captures both local and global features while targeting blur-related frequencies, achieving competitive results with state-of-the-art methods at faster inference speeds. Through extensive experiments on benchmark datasets and human evaluations, we demonstrate that we technique produces high-quality results across diverse image types, making it well-suited for real-world applications.

Sharp Image
CodeStream: A Dataset of Iterative Programming Submissions with Sequential Verdict Traces and Attempt Histories
Nazira Jesmin Lina, Syed Mumtahin Mahmud, Mahmudul Hasan, Md Fahim Arefin, Redwan Ahmed Rizvee, Md Mahmudur Rahman, Md Mosaddek Khan
Data in Brief, Volume 66, Elsevier, 2026

We present CodeStream, a dataset capturing how novice programmers converge on correct solutions across repeated attempts. Collected from proctored sessions on a custom assessment platform, it contains 5,482 submissions from 202 undergraduate participants across 46 problems in C, C++, and Java. Unlike datasets logging only final outcomes, every record preserves the attempt index, full source code, and the ordered per-test-case verdict trace, linked to problem statements and their evaluation test cases. The dataset supports research in educational data mining, learning analytics, automated feedback, and code analysis, and is available on Mendeley Data.

Work Experience

University of Dhaka

Research Assistant, Cognitive Agents and Interaction Lab, University of Dhaka
July 2025 - Present
Advised by Dr. Md Mosaddek Khan
Working on projects related to image deblurring, medical imaging, LLM, and Graph Neural Network.

University of Dhaka

Team Lead (Python Dev), Turing Enterprises, Inc.
October 2024 - June 2025 (Remote)
Led a team of 5 in LLM data creation. Designed and implemented data pipelines and preprocessing workflows to generate specialized training data for Agentic AI behaviors, includ‑ ing multistep reasoning and decision making scenarios. Created specialized training datasets for agentless AI approaches, developing data generation methodologies that enable direct task completion without explicit agent frameworks or intermediary reasoning steps.

University of Dhaka

Junior Data Analyst, One Data Labs
May 2024 - Septemeber 2024
Actively involved in the meticulous examination and transformation of unprocessed data to extract valuable insights, streamline operations, and support strategic initiatives, ensuring accuracy & integrity.

International Programs

IIT Bombay Summer School Program 2024 International Summer School

Indian Institute of Technology, Bombay

Mumbai, Maharashtra, India

Course: Quantum Computing for Machine Learning and Optimization

Participated in the onsite program. Selected as one of only 4 students from Bangladesh for this prestigious international summer school.

June 2024

Miscellanea

Programming Contests

BUET Inter-University Programming Contest (2022) - Team: DU_Bottom_Frag
SEC Inter-University Programming Contest (2022) - Team: DU_Bottom_Frag
Meta Hacker Cup (2022)
Meta Hacker Cup (2021)
Code Samurai Inter University Hackathon (2022) - Team: DU_Restoring_Over_Non_Restoring