Curriculum Vitae
Ashmit Mukherjee
Applied AI, computational experimentation, and data-driven research at New York University Abu Dhabi. [email protected]
Education
New York University Abu Dhabi
Expected May 2027B.S. in Computer Science. Studied at NYU Abu Dhabi, NYU Paris, and NYU New York.
- Relevant coursework: Natural Language Processing; Principles of Data Science; Statistics for Social & Behavioral Sciences; Data Structures & Algorithms; Computer Systems Architecture; Operating Systems.
Research Interests
I build and evaluate AI systems using data, statistical inference, machine learning, experimental design, and research engineering. My projects span the full empirical workflow: constructing datasets, developing models and systems, designing interventions, evaluating results, and connecting technical evidence to human or social outcomes. I currently apply these methods to human-AI collaboration, parameter-efficient model adaptation, and multilingual NLP.
Publications and Working Papers
Benjamin Rosche, Ashmit Mukherjee, and Hanan Salam. “AI and Human Collaboration in Social Dilemmas: Strategic Interpretation, Normative Interpretation, and Collective Action in Shared Supply Chains.” Working paper, June 2026.
Ashmit Mukherjee and collaborators. “Bio-Informed LoRA for Signal Peptide Prediction.” Co-first author. Under review at an EMNLP 2026 workshop, 2026.
Research Experience
Research Assistant: Human-AI Collaboration
Feb 2026 to PresentNYU Abu Dhabi
- Co-developing an online experiment on how AI-mediated interpretation and coordination affect collective action in shared global supply chains, with Benjamin Rosche and Hanan Salam.
- Designed a repeated sourcing game in which five human buyers choose fair pay, auditing, information-sharing links, communication, and threshold-dependent remediation around one supplier governed by a fixed decision rule.
- Developed a group-level, four-arm treatment design comparing structured information with strategic interpretation AI, normative interpretation AI, and collective action AI.
- Specified behavioral outcomes and mechanism measures for cooperation, responsibility attribution, behavioral homogenization, coalition formation, and persistence following AI withdrawal.
Research Assistant: eBRAIN Lab
Jan 2026 to PresentNYU Abu Dhabi
- Co-developed Bio-Informed LoRA, a parameter-efficient fine-tuning method that injects residue-level biological priors (BLOSUM62, hydrophobicity, Grantham) into LoRA adapters for ESM-2 protein language models on signal peptide prediction.
- Showed BLOSUM62-guided LoRA matches or exceeds full fine-tuning on SignalP6 benchmarks across MCC, cleavage-site exactness, and residue-level F1, while training only ~3.6% of parameters and using ~43% less peak GPU memory at the 3B-parameter scale.
- Designed multi-seed evaluation protocols and a one-factor sensitivity screen showing that several single-seed improvements did not replicate.
- Manuscript under review (see Publications).
Research Projects
Hinglish Named Entity Recognition Benchmark
Fall 2024- Fine-tuned multilingual transformer models, mBERT and XLM-RoBERTa, on the COMI-LINGUA dataset for Hindi-English code-mixed NER.
- Compared against zero-shot GPT-4o and Claude 3.5 Sonnet under a common evaluation protocol; fine-tuned XLM-RoBERTa achieved 78% entity-level F1 versus 76% for GPT-4o.
- Analyzed trade-offs between model scale, task-specific fine-tuning, generalization, and computational cost.
Data Science Salary Prediction Platform (Methodological Study)
Fall 2025- Benchmarked 15+ regression models using AutoML (PyCaret) on a dataset of 3,000+ job postings.
- Used SHAP to interpret associations between model predictions and features including geography, firm size, and skill requirements.
- Tracked experiments with MLflow; reported MAE, RMSE, and R² across conditions.
Community-Level Crime Modeling
Spring 2025- Modeled statistical associations between socioeconomic indicators and violent crime rates across 1,994 communities without making causal claims.
- Evaluated predictive performance using cross-validation (R² = 0.85) and built interactive visualizations for interpretation.
CAMP: Campus Asset Management Platform
Fall 2025- Contributed to a team-built system for resource allocation and scheduling under shared constraints, including role-based access and conflict resolution.
- Developed RESTful APIs and automated CI/CD pipelines using Docker and GitHub Actions for reproducible deployment.
Industry Experience
Machine Learning Engineer Intern
Jun to Aug 2025Zeek
- Developed and evaluated segmentation models on noisy real-world data, iterating on preprocessing and model selection for production reliability.
- Applied model interpretability techniques to study robustness and bias sensitivity, helping the team decide which model variants to deploy.
- Contributed to the team’s evaluation workflow by adding standardized metrics and validation procedures to the model development cycle.
Technical Skills
- Languages
- Python, C++, C#, JavaScript
- ML & NLP
- PyTorch, Transformers (Hugging Face), Scikit-Learn, XGBoost, PyCaret, SHAP
- Computational Modeling
- NumPy, NetworkX, game-theoretic simulation
- Data Analysis
- Pandas, SciPy, Matplotlib, Seaborn, Plotly
- Experimental Methods
- Experimental design, regression, cross-validation, multi-seed evaluation, sensitivity analysis
- Research Engineering
- MLflow, Git, Docker, GitHub Actions, REST APIs, automated testing, MongoDB
- Applications
- Streamlit, React, Node.js, LaTeX