Curriculum Vitae

Ashmit Mukherjee

Applied AI, computational experimentation, and data-driven research at New York University Abu Dhabi. [email protected]

Education

New York University Abu Dhabi

Expected May 2027

B.S. in Computer Science. Studied at NYU Abu Dhabi, NYU Paris, and NYU New York.

  • Relevant coursework: Natural Language Processing; Principles of Data Science; Statistics for Social & Behavioral Sciences; Data Structures & Algorithms; Computer Systems Architecture; Operating Systems.

Research Interests

I build and evaluate AI systems using data, statistical inference, machine learning, experimental design, and research engineering. My projects span the full empirical workflow: constructing datasets, developing models and systems, designing interventions, evaluating results, and connecting technical evidence to human or social outcomes. I currently apply these methods to human-AI collaboration, parameter-efficient model adaptation, and multilingual NLP.

Publications and Working Papers

Benjamin Rosche, Ashmit Mukherjee, and Hanan Salam. “AI and Human Collaboration in Social Dilemmas: Strategic Interpretation, Normative Interpretation, and Collective Action in Shared Supply Chains.” Working paper, June 2026.

Ashmit Mukherjee and collaborators. “Bio-Informed LoRA for Signal Peptide Prediction.” Co-first author. Under review at an EMNLP 2026 workshop, 2026.

Research Experience

Research Assistant: Human-AI Collaboration

Feb 2026 to Present

NYU Abu Dhabi

  • Co-developing an online experiment on how AI-mediated interpretation and coordination affect collective action in shared global supply chains, with Benjamin Rosche and Hanan Salam.
  • Designed a repeated sourcing game in which five human buyers choose fair pay, auditing, information-sharing links, communication, and threshold-dependent remediation around one supplier governed by a fixed decision rule.
  • Developed a group-level, four-arm treatment design comparing structured information with strategic interpretation AI, normative interpretation AI, and collective action AI.
  • Specified behavioral outcomes and mechanism measures for cooperation, responsibility attribution, behavioral homogenization, coalition formation, and persistence following AI withdrawal.

Research Assistant: eBRAIN Lab

Jan 2026 to Present

NYU Abu Dhabi

  • Co-developed Bio-Informed LoRA, a parameter-efficient fine-tuning method that injects residue-level biological priors (BLOSUM62, hydrophobicity, Grantham) into LoRA adapters for ESM-2 protein language models on signal peptide prediction.
  • Showed BLOSUM62-guided LoRA matches or exceeds full fine-tuning on SignalP6 benchmarks across MCC, cleavage-site exactness, and residue-level F1, while training only ~3.6% of parameters and using ~43% less peak GPU memory at the 3B-parameter scale.
  • Designed multi-seed evaluation protocols and a one-factor sensitivity screen showing that several single-seed improvements did not replicate.
  • Manuscript under review (see Publications).

Research Projects

Hinglish Named Entity Recognition Benchmark

Fall 2024
  • Fine-tuned multilingual transformer models, mBERT and XLM-RoBERTa, on the COMI-LINGUA dataset for Hindi-English code-mixed NER.
  • Compared against zero-shot GPT-4o and Claude 3.5 Sonnet under a common evaluation protocol; fine-tuned XLM-RoBERTa achieved 78% entity-level F1 versus 76% for GPT-4o.
  • Analyzed trade-offs between model scale, task-specific fine-tuning, generalization, and computational cost.

Data Science Salary Prediction Platform (Methodological Study)

Fall 2025
  • Benchmarked 15+ regression models using AutoML (PyCaret) on a dataset of 3,000+ job postings.
  • Used SHAP to interpret associations between model predictions and features including geography, firm size, and skill requirements.
  • Tracked experiments with MLflow; reported MAE, RMSE, and R² across conditions.

Community-Level Crime Modeling

Spring 2025
  • Modeled statistical associations between socioeconomic indicators and violent crime rates across 1,994 communities without making causal claims.
  • Evaluated predictive performance using cross-validation (R² = 0.85) and built interactive visualizations for interpretation.

CAMP: Campus Asset Management Platform

Fall 2025
  • Contributed to a team-built system for resource allocation and scheduling under shared constraints, including role-based access and conflict resolution.
  • Developed RESTful APIs and automated CI/CD pipelines using Docker and GitHub Actions for reproducible deployment.

Industry Experience

Machine Learning Engineer Intern

Jun to Aug 2025

Zeek

  • Developed and evaluated segmentation models on noisy real-world data, iterating on preprocessing and model selection for production reliability.
  • Applied model interpretability techniques to study robustness and bias sensitivity, helping the team decide which model variants to deploy.
  • Contributed to the team’s evaluation workflow by adding standardized metrics and validation procedures to the model development cycle.

Technical Skills

Languages
Python, C++, C#, JavaScript
ML & NLP
PyTorch, Transformers (Hugging Face), Scikit-Learn, XGBoost, PyCaret, SHAP
Computational Modeling
NumPy, NetworkX, game-theoretic simulation
Data Analysis
Pandas, SciPy, Matplotlib, Seaborn, Plotly
Experimental Methods
Experimental design, regression, cross-validation, multi-seed evaluation, sensitivity analysis
Research Engineering
MLflow, Git, Docker, GitHub Actions, REST APIs, automated testing, MongoDB
Applications
Streamlit, React, Node.js, LaTeX