Click here for a printable version!


David Nicholson

dnicholson329@gmail.com • 412-607-6313 • Lansdowne, PA (Open to Remote)

Summary

Data Scientist experienced in architecting machine learning workflows, fine-tuning LLMs, and managing multi-tenant database systems. Adept at bridging technical data engineering with business decision-making through interactive dashboards, automated metadata pipelines, and production-ready NLP models.

Skills & Proficiency

  • Machine Learning & NLP: PyTorch, TensorFlow, Scikit-Learn, Transformers, Large Language Models (Claude, Gemini, Ollama), Clustering (HDBSCAN, K-Means), Dimensionality Reduction (UMAP), Document Embeddings, Knowledge Graphs, Entity Resolution, Text Mining and Extraction, Automatic Metadata Extraction, Weak Supervision (Snorkel)
  • Data Engineering & Systems: Google Bigquery, PostgreSQL, ETL/ELT Pipelines, Parallel Processing (Dask), Linux Server Maintenance
  • Dashboards & Visualization: Plotly/Dash, Seaborn/Matplotlib, Web Applications (Flask/FastAPI)
  • Core Infrastructure: Python, SQL, Git, Docker, CI/CD, AWS/GCP

Professional Experience

Data Scientist
Digital Science & Research Solutions, Ltd.
June 2022 - Present

  • LLM & Automated Annotation: Constructed an LLM-backed document metadata and labeling pipeline to automatically fetch, categorize and summarize clusters across multiple document types for United States Federal Research Funding Agencies and Institutes.
  • Vector Databases & Scalable Search: Conducted benchmarking across enterprise vector databases to optimize query latency and storage efficiency for 10M+ high-dimensional document vectors.
  • ETL & Big Data Engineering: Engineered automated ETL pipelines to extract external unstructured web/literature data alongside secure ingestion and processing of HHS-affiliated client’s data using Python and SQL. Also, maintained and modified other ETL pipelines for processing 100GB+ of unstructured literature data.
  • Dashboards & Reporting: Built interactive dashboards (Plotly/Streamlit) for internal use to explore dynamic topic clusters, research trends, and growth metrics in real time.

Graduate Researcher Scientist
University of Pennsylvania
August 2016 - June 2022

  • Large-Scale NLP & Computational Modeling: Designed and deployed unsupervised NLP frameworks (Word2Vec, FastText) to analyze over 20,000 temporal semantic shifts across millions of unstructured scientific documents.
  • Pipeline Engineering & Big Data: Built automated, parallelized Python pipelines to ingest, clean, and analyze multi-gigabyte genomic and textual datasets, reducing data processing bottlenecks by over 40%.
  • Statistical Modeling & Algorithm Optimization: Applied high-dimensional dimensionality reduction (UMAP) and clustering algorithms (K-Means) to extract hidden patterns in complex biological networks.
  • Cross-Functional Collaboration & Technical Communication: Collaborated with cross-departmental teams of bioinformaticians, software engineers, and domain experts; presented complex quantitative findings to non-technical stakeholders and external partners.
  • Software Development & Maintainability: Maintained open-source codebase repositories on GitHub using strict version control (Git/GitHub), unit testing, and Docker containerization.

Publications

Education

Doctor of Philosophy (Ph.D.), Genomics and Computational Biology; University of Pennsylvania (Philadelphia, PA)

Postbaccalaureate Program (Penn Prep); University of Pennsylvania (Philadelphia, PA)

Bachelor of Science, Computer Science; University of Maryland Baltimore County (Baltimore, MD)