Bioinformatics PhD candidate studying gene duplication and lineage-specific gene evolution in Populus through comparative transcriptomics and phylogenetics, from genome-scale pipelines down to transgenic validation in the greenhouse and field. Recently stress-testing genomic foundation models proposed for DNA synthesis biosecurity screening: parameter-efficient fine-tuning on DNA and RNA language models, plus causal interpretability to establish what these models actually read and whether that signal can be defended.
Available Fall 2026. Based in Athens, Georgia. chen.hsieh.uga@gmail.com · GitHub · LinkedIn · Google Scholar
Education
| University of Georgia — PhD Candidate, Bioinformatics | Expected Dec 2026 |
| National Taiwan University — MSc, Horticultural Crops Science | Jun 2019 |
| National Taiwan University — BS, Agriculture | Jun 2017 |
Technical skills
Programming: Python · R · Bash/Shell · JavaScript · C/C++
Deep learning & foundation models: PyTorch · Hugging Face · DNA and RNA language models across nucleotide-, k-mer- and codon-tokenized families and transformer and state-space architectures · LoRA · parameter-efficient fine-tuning · embedding extraction, pooling, and frozen probing
Interpretability & robustness: per-layer linear probing · tuned lens · activation patching and direct feature erasure · in-silico mutagenesis · per-position importance attribution · decision-depth analysis · black-box and white-box adversarial attacks and training · matched random-edit floors · benchmark-validity auditing
Genomics & classical ML: RNA-seq pipelines (STAR, HISAT2, DESeq2, edgeR) · 16S metagenomics (DADA2, QIIME2, SILVA, α/β diversity, MaAsLin2) · comparative transcriptomics · phylogenetics · genome-wide annotation · CRISPR gRNA design · XGBoost and SHAP
Engineering & deployment: Git/GitHub · Docker · CI/CD (GitHub Actions) · Python packaging (PyPI) · REST API · Snakemake · HPC/SLURM · agentic research workflows · Streamlit / R Shiny
Research experience
PhD Dissertation Research — Tsai Lab
Institute of Bioinformatics, University of Georgia · 2020 – present
- Dissertation on lineage-specific genes and the gene-duplication landscape in Populus, using expert-annotated and published case studies as a ground-truth reference set to improve automated analysis pipelines.
- Built the lineage-specific gene identification pipeline (phylostratigraphy with GenEra plus synteny, homology-detection-failure, and outgroup quality control), then established length as the dominant confounder and length-matched every downstream comparison.
- Characterized the resulting genes across sequence, structure, and expression: neighbor-proximity as a predictor of transcription, high hydropathy with intrinsic disorder, Foldseek-based structural homology recovery beyond BLAST, and a 652-sample, 11-tissue expression atlas showing bark and xylem enrichment and stress inducibility.
- Carried candidates from ~65,000 gene models to six transgenic constructs through greenhouse drought trials and a field nursery planting, with COLD-REGULATED 15A as a positive control.
- Reprocessed large-scale community RNA-seq data (FASTQ to differential expression) on reproducible, HPC-scalable pipelines; built genome-wide protein structure annotation workflows.
- Modeled ohnolog expression divergence with XGBoost and SHAP under leakage-aware validation (chromosome-held-out cross-validation, permuted-label controls).
- Evaluated seven publicly available DNA foundation models, spanning several architectures and tokenization schemes, on Populus promoter sequences — diagnosing embedding quality and establishing architecture-specific pooling.
- Contributed to VariantDB v3.2 (CRISPR gRNA verification tool) and coexpression analysis of a perennial-specific sulfate transporter subgroup associated with lignification; both published in Tree Physiology (2025).
What Genomic Foundation Models Read: Composition vs. Arrangement
Algoverse AI Research Program, team research project · 2026
- Designed a composition-exact counterfactual benchmark to separate what genomic language models read from what a lookup table could: a few thousand coding sequences across several taxa from a single corpus, eight synonymous-rewrite conditions, and a per-sequence assertion that codon composition is unchanged rather than assumed.
- Scored it two independent ways — task-coupled (per-layer embeddings through one fixed linear head, balanced-accuracy drop, stratified paired bootstrap) and label-free (the model’s own masked-codon distribution, no head, no labels, no split) — and anchored both to model-free reference levels (gradient boosting on codon frequency, a summary-statistic classifier).
- Established that arrangement carries real but small signal relative to composition, and that sensitivity varies widely across codon- and nucleotide-tokenized models spanning more than an order of magnitude in size, with no monotone trend.
- Showed a linear probe and a direct causal intervention disagree about which feature drives the decision, using tuned lens, per-layer probing, and direct feature erasure — decodability is not use.
- Evaluated LoRA, head-only, and full fine-tuning as defenses, and characterized transfer between models. Related manuscript submitted to a workshop in August 2026; not yet peer-reviewed. Application is DNA synthesis biosecurity screening; specifics available on request rather than published ahead of review.
MSc Research
Department of Horticulture, National Taiwan University · 2017 – 2019
- Comparative transcriptomics of mycorrhiza-enhanced salt tolerance in rice (Frontiers in Plant Science, 2022, co-first author).
- Automated HRM genotyping pipelines for a peach chilling requirement study (IJMS, 2020).
Selected leadership & projects
Technology Consultant & IT Group Convener
Taiwanese Young Researcher Association (Project TYRA) · Apr 2022 – present
- Built a platform powering a mentor–mentee matching program that grew from 274 to 700+ annual participants across 4 cohorts and facilitated 1,876 researcher connections; organization later partnered with Fulbright Taiwan and government education agencies.
- Deployed an automated event pipeline (Eventbrite API plus scheduling automation) for the weekly seminar series.
Director of Marketing & Communication
PhD Consulting Club, University of Georgia · Jul – Dec 2025
- Tripled club membership across multiple departments; facilitated weekly structured case practice sessions.
Team Leader
Hackathons & startup competitions · 2017 – present
- Led cross-functional teams of up to 5 at HudsonAlpha HATCH (2024, 2025) and the UGA-Bayer Hackathon (2025); delivered working prototypes under 8- to 28-hour constraints, including a plant-based diet planner for astronauts and a corn yield predictor (Snakemake, PyTorch, OpenAI API, Streamlit).
- Built the first Mandarin-language chatbot for plant disease diagnosis (Open Data Innovative Application Contest, Taiwan, 2017); pitched to venture capital panels and secured an NTD$410,000 award.
Publications
Five peer-reviewed publications. Full citations on the publications page.
- Surber et al. (2025), Tree Physiology — sulfate transporter phylogeny, perennial-specific subgroup
- Zhou et al. (2025), Tree Physiology — Populus VariantDB v3.2
- Tuma et al. (2024), Tree Physiology — tonoplast sucrose transport in coppiced poplar
- Hsieh et al. (2022), Frontiers in Plant Science — mycorrhiza-enhanced salt tolerance in rice (co-first author)
- Chou et al. (2020), IJMS — HRM genotyping toolkit for peach chilling requirement
Selected presentations
Talks. Stress-testing a codon language model for DNA synthesis screening, Institute of Bioinformatics student seminar, UGA (2026) · Lineage-specific genes in a hybrid poplar genome, Tsai Lab / Institute of Bioinformatics, UGA (2026) · Investigating de novo gene birth in Populus, SMBE Satellite Meeting on De Novo Gene Birth, Texas A&M University (2023).
Posters. Characterization of stress-responsive lineage-specific genes in Populus, IUFRO Tree Biotechnology Conference (2024) · Searching for orphan genes in Populus, ASPB Worldwide Summit (2021) · Comparative transcriptome analysis of gibberellin-induced sex determination in bitter gourd, XXX International Horticultural Congress, Istanbul (2018).
Selected open-source tools
Research. slurm-receipt (PyPI): CLI that turns SLURM job history into a compute-cost, energy, and cloud-equivalent report · sapelo2-boilerplate: documented sbatch recipes for eight genomics tools, Snakemake pipelines, and a Claude Code ruleset for GPU/HPC jobs · agentic-research-toolkit: portable agentic research workflows (SKILL.md) written for discovery over confirmation · Biosecurity Atlas: knowledge graph of the biosafety-and-AI research funding landscape aggregated from five public funding and publication APIs.
Creative coding. dna2oiia (PyPI, meme-inspired DNA-to-audio sonification) · bioLOLPython (PyPI, sequence analysis with internet-slang dialects) · Glasshouse, a research greenhouse rendered as a game of Snake · UGA Grad Survivor, a Reigns-style PhD-survival game · slang-capsule, a multilingual timeline of internet slang queryable by year, language, and platform.
More detail on the projects page.
Teaching, mentoring & outreach
Mentored undergraduate researchers in bioinformatics workflows and phenotyping (UGA, 2021–2025) and in genome-wide sequence analysis (NTU, 2017–2018), where the mentee won first prize at the departmental poster competition. Guest lecturer, Methods in Horticultural Research (IV), National Taiwan University: taught RNA-seq analysis to an audience with no computational background; pre-lecture content doubled expected attendance. Elected BIGSA Representative (2022–2023) and organizer of the Institute of Bioinformatics seminar series, hosting Josh Starmer (StatQuest), Robert Edgar (MUSCLE), Ben Langmead (Bowtie/Bowtie2), and Jeffrey Perkel (Nature Technology Editor).
More on the teaching and community page.
Awards & certificates
Government Fellowship for Studying Abroad (est. $150K), Ministry of Education, Taiwan (2019) · Summer Research & Communication Grants, UGA (2021, 2022) · Sci4Pol Certificate, National Science Policy Network (2025).