Research (Professional) Escobar and Hu Lab in Pittsburgh, Pennsylvania at University of Pittsburgh
Explore Related Opportunities
Job Description
The Research Assistant will play a core technical role in the PROXIMO project, an interdisciplinary clinical AI research initiative focused on privacy-preserving data preparation, responsible development, and evaluation of large language model (LLM)–based mental health support chatbots.
The assistant will lead and optimize technical workflows involving de-identified conversational data, driving the development of robust data pipelines, quality-control protocols, and scalable research architectures. This position is highly suited for a Master’s or Ph.D. student with a strong background in artificial intelligence, natural language processing (NLP), and large language models, looking to apply advanced technical skills to complex healthcare technology and data privacy challenges.
Specific responsibilities include designing, optimizing, and maintaining advanced Python pipelines for conversational text processing; implementing and refining privacy-preserving workflows using state-of-the-art anonymization tools (e.g., Microsoft Presidio); constructing structured, high-quality datasets for LLM fine-tuning and evaluation; and establishing rigorous, reproducible research frameworks for the PROXIMO team.
The assistant will actively drive the development and quantitative/qualitative evaluation of LLM-based mental health support chatbots. Activities will include architecting multi-turn conversational data for model training/evaluation, designing sophisticated prompt engineering and evaluation metrics, evaluating model outputs for safety, empathy, tone, and clinical escalation protocols, and translating complex technical findings into actionable insights for an interdisciplinary research team.
While the assistant will inherit existing code and documentation, they are expected to bring the technical maturity to independently audit, refactor, and scale these workflows, taking ownership of the technical infrastructure. This role involves close collaboration with domain experts in psychiatry, psychology, computer science, and engineering. Exceptional attention to data privacy, reproducible coding practices, and responsible AI principles is essential.
Required Qualifications
- Current Master's or Ph.D.student in computer science, computer engineering, electrical engineering, data science, information science, statistics, or a closely related highly technical field.
- Advanced programming proficiency in Python, with a strong grasp of software engineering best practices, version control (Git/GitHub), and code reproducibility.
- Demonstrated experience building, training, or evaluating Natural Language Processing (NLP) pipelines and Large Language Models (LLMs) through research, industry internships, or substantial thesis projects.
- Strong hands-on experience with core data science and deep learning libraries (e.g., PyTorch, Hugging Face Transformers, pandas, NumPy).
- Ability to independently architect, test, and troubleshoot complex data processing and machine learning workflows with minimal micromanagement.
- Solid understanding of responsible AI principles, data privacy constraints, and the ethical implications of deploying AI in healthcare settings.
- Excellent technical writing and communication skills, with the ability to distill complex algorithmic concepts and data analyses for interdisciplinary collaborators (e.g., clinical psychologists and psychiatrists).
- High level of meticulousness and attention to detail, particularly regarding the handling and validation of sensitive clinical-adjacent data.
Preferred Qualifications
- Current Ph.D. student or a Master's student with a thesis/research focus on NLP, Generative AI, or Digital Health.
- Prior experience fine-tuning open-weights LLMs (e.g., Llama, Mistral), implementing Retrieval-Augmented Generation (RAG) architectures, or developing agentic workflows.
- Experience operating in high-performance computing (HPC) or cloud-based research computing environments.
- Familiarity with state-of-the-art de-identification, data anonymization frameworks, or privacy-preserving machine learning techniques.
- Prior experience working with raw, unstructured conversational, chat-style, or clinical text datasets.
- A track record of academic publications, technical reports, or a strong GitHub portfolio demonstrating relevant NLP/AI capabilities.
- Experience mentoring or guiding junior researchers/undergraduate students in a lab setting.
- Familiarity with human-centered AI evaluation methods, chatbot safety guardrails, or psychological/mental health domain knowledge.
Example Tasks
Depending on project needs and the student's research focus, tasks will include:
- Pipeline Architecture: Designing, refactoring, and scaling Python-based NLP pipelines to process large volumes of conversational data efficiently.
- Privacy Infrastructure: Enhancing and automating de-identification algorithms (e.g., customizing Microsoft Presidio) to ensure rigorous privacy preservation of sensitive text.
- LLM Development: Deploying, testing, and fine-tuning open-source LLMs within secure, approved computing environments to simulate mental health chatbot interactions.
- Evaluation Design: Formulating and executing robust evaluation frameworks (both automated metrics and human-in-the-loop rubrics) to assess chatbot safety, clinical adherence, and empathy.
- Data Engineering: Preparing highly structured, quality-assured datasets for downstream model training and clinical analysis.
- Technical Leadership: Establishing standardized Markdown documentation, GitHub CI/CD practices (if applicable), and reproducible workflows for the entire PROXIMO engineering team.
- Research Dissemination: Conducting sophisticated data analyses and co-authoring technical summaries, data visualizations, and potentially academic manuscripts for interdisciplinary dissemination.
Master's degree or PhD Student
The University of Pittsburgh is an equal opportunity employer / disability / veteran.
Department: Electrical and Computer Engineering
Campus: Pittsburgh
Minimum Education Level Required: Master's Degree
Minimum Years of Experience Required: 2
Work Schedule: Monday - Friday, 8:30am - 5pm
Work Arrangement: On-Campus: Teams that work on campus, in an office, or in a lab.
Requested Pay Rate: 30.00
Visa Sponsorship Provided:
Background Check: For position finalists, employment with the University will require successful completion of a background check
Child Protection Clearances: Not Applicable
Required Documents: Resume, Cover Letter