Roman Cheplyaka

the bioinformatics chat

Science EN ↓ 70 episodes

A podcast about computational biology, bioinformatics, and next generation sequencing.

Author

Roman Cheplyaka

Category

Science

Podcast website

bioinformatics.chat

Latest episode

Dec 21, 2023

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

#45 Genome assembly and Canu with Sergey Koren and Sergey Nurk 20.05.2020

In this episode, Sergey Nurk and Sergey Koren from the NIH share their thoughts on genome assembly. The two Sergeys tell the stories behind their amazing careers as well as behind some of the best known genome assemblers: Celera assembler, Canu, and SPAdes. Links: Canu on GitHub SPAdes on GitHub

#44 DNA tagging and Porcupine with Kathryn Doroschak 29.04.2020

Porcupine is a molecular tagging system—a way to tag physical objects with pieces of DNA called molecular bits , or molbits for short. These DNA tags then can be rapidly sequenced on an Oxford Nanopore MinION device without any need for library preparation. In this episode, Katie Doroschak explains how Porcupine works—how molbits are designed and prepared, and how they are directly recognized by t...

#43 Generalized PCA for single-cell data with William Townes 27.03.2020

Will Townes proposes a new, simpler way to analyze scRNA-seq data with unique molecular identifiers (UMIs). Observing that such data is not zero-inflated, Will has designed a PCA-like procedure inspired by generalized linear models (GLMs) that, unlike the standard PCA, takes into account statistical properties of the data and avoids spurious correlations (such as one or more of the top principal c...

#42 Spectrum-preserving string sets and simplitigs with Amatur Rahman and Karel Břinda 28.02.2020

In this episode, we hear from Amatur Rahman and Karel Břinda , who independently of one another released preprints on the same concept, called simplitigs or spectrum-preserving string sets. Simplitigs offer a way to efficiently store and query large sets of k-mers—or, equivalently, large de Bruijn graphs. Links: Simplitigs as an efficient and scalable representation of de Bruijn graphs (Karel Břin...

#41 Epidemic models with Kris Parag 27.01.2020

Kris Parag is here to teach us about the mathematical modeling of infectious disease epidemics. We discuss the SIR model, the renewal models, and how insights from information theory can help us predict where an epidemic is going. Links: Optimising Renewal Models for Real-Time Epidemic Prediction and Estimation (KV Parag, CA Donnelly) Adaptive Estimation for Epidemic Renewal and Phylogenetic Skyli...

#40 Plasmid classification and binning with Sergio Arredondo-Alonso and Anita Schürch 30.12.2019

Does a given bacterial gene live on a plasmid or the chromosome? What other genes live on the same plasmid? In this episode, we hear from Sergio Arredondo-Alonso and Anita Schürch , whose projects mlplasmids and gplas answer these types of questions. Links: mlplasmids: a user-friendly tool to predict plasmid- and chromosome-derived sequences for single species (Sergio Arredondo-Alonso, Malbert R....

#39 Amplicon sequence variants and bias with Benjamin Callahan 29.11.2019

In this episode, Benjamin Callahan talks about some of the issues faced by microbiologists when conducting amplicon sequencing and metagenomic studies. The two main themes are: Why one should probably avoid using OTUs (operational taxonomic units) and use exact sequence variants (also called amplicon sequence variants, or ASVs), and how DADA2 manages to deduce the exact sequences present in the sa...

#38 Issues in legacy genomes with Luke Anderson-Trocmé 22.10.2019

In this episode, Luke Anderson-Trocmé talks about his findings from the 1000 Genomes Project. Namely, the early sequenced genomes sometimes contain specific mutational signatures that haven’t been replicated from other sources and can be found via their association with lower base quality scores. Listen to Luke telling the story of how he stumbled upon and investigated these fake variants and what...

#37 Causality and potential outcomes with Irineo Cabreros 27.09.2019

In this episode, I talk with Irineo Cabreros about causality. We discuss why causality matters, what does and does not imply causality, and two different mathematical formalizations of causality: potential outcomes and directed acyclic graphs (DAGs). Causal models are usually considered external to and separate from statistical models, whereas Irineo’s new paper shows how causality can be viewed a...

#36 scVI with Romain Lopez and Gabriel Misrachi 30.08.2019

In this episode, we hear from Romain Lopez and Gabriel Misrachi about scVI—Single-cell Variational Inference. scVI is a probabilistic model for single-cell gene expression data that combines a hierarchical Bayesian model with deep neural networks encoding the conditional distributions. scVI scales to over one million cells and can be used for scRNA-seq normalization and batch effect removal, dimen...

#35 The role of the DNA shape in transcription factor binding with Hassan Samee 26.07.2019

Even though the double-stranded DNA has the famous regular helical shape, there are small variations in the geometry of the helix depending on what exact nucleotides its made of at that position. In this episode of the bioinformatics chat, Hassan Samee talks about the role the DNA shape plays in recognition of the DNA by DNA-binding proteins, such as transcription factors. Hassan also explains how...

#34 Power laws and T-cell receptors with Kristina Grigaityte 29.06.2019

An αβ T-cell receptor is composed of two highly variable protein chains, the α chain and the β chain. However, based only on bulk DNA or RNA sequencing it is impossible to determine which of the α chain and β chain sequences were paired in the same receptor. In this episode, Kristina Grigaityte talks about her analysis of 200,000 paired αβ sequences, which have been obtained by targeted single-cel...

#33 Genome assembly from long reads and Flye with Mikhail Kolmogorov 31.05.2019

Modern genome assembly projects are often based on long reads in an attempt to bridge longer repeats. However, due to the higher error rate of the current long read sequencers, assemblers based on de Bruijn graphs do not work well in this setting, and the approaches that do work are slower. In this episode, Mikhail Kolmogorov from Pavel Pevzner’s lab joins us to talk about some of the ideas develo...

#32 Deep tensor factorization and a pitfall for machine learning methods with Jacob Schreiber 29.04.2019

In this episode, we hear from Jacob Schreiber about his algorithm, Avocado. Avocado uses deep tensor factorization to break a three-dimensional tensor of epigenomic data into three orthogonal dimensions corresponding to cell types, assay types, and genomic loci. Avocado can extract a low-dimensional, information-rich latent representation from the wealth of experimental data from projects like the...

#31 Bioinformatics Contest 2019 with Alexey Sergushichev and Gennady Korotkevich 24.03.2019

The third Bioinformatics Contest took place in February 2019. Alexey Sergushichev , one of the organizers of the contest, and Gennady Korotkevich , the 1st prize winner, join me to discuss this year’s problems. Timestamps and links for the individual problems: Qualification round 00:07:14 Bee Population 00:14:12 Sequencing Errors 00:30:20 Transposable Elements Final round 00:41:35 Cancer and Chrom...

#30 Bayesian inference of chromatin structure from Hi-C data with Simeon Carstens 27.02.2019

Hi-C is a sequencing-based assay that provides information about the 3-dimensional organization of the genome. In this episode, Simeon Carstens explains how he applied the Inferential Structure Determination (ISD) framework to build a 3D model of chromatin and fit that model to Hi-C data using Hamiltonian Monte Carlo and Gibbs sampling. Links: Bayesian inference of chromatin structure ensembles fr...

#29 Haplotype-aware genotyping from long reads with Trevor Pesout 27.01.2019

Long read sequencing technologies, such as Oxford Nanopore and PacBio, produce reads from thousands to a million base pairs in length, at the cost of the increased error rate. Trevor Pesout describes how he and his colleagues leverage long reads for simultaneous variant calling/genotyping and phasing. This is possible thanks to a clever use of a hidden Markov model, and two different algorithms ba...

#28 Space-efficient variable-order Markov models with Fabio Cunial 28.12.2018

This time you’ll hear from Fabio Cunial on the topic of Markov models and space-efficient data structures. First we recall what a Markov model is and why variable-order Markov models are an improvement over the standard, fixed-order models. Next we discuss the various data structures and indexes that allowed Fabio and his collaborators to represent these models in a very small space while still ke...

#27 Classification of CRISPR-induced mutations and CRISPRpic with HoJoon Lee and Seung Woo Cho 29.11.2018

In this episode, HoJoon Lee and Seung Woo Cho explain how to perform a CRISPR experiment and how to analyze its results. HoJoon and Seung Woo developed an algorithm that analyzes sequenced amplicons containing the CRISPR-induced double-strand break site and figures out what exactly happened there (e.g. a deletion, insertion, substitution etc.) Links: CRISPRpic: Fast and precise analysis for CRISPR...

#26 Feature selection, Relief and STIR with Trang Lê 27.10.2018

Relief is a statistical method to perform feature selection. It could be used, for instance, to find genomic loci that correlate with a trait or genes whose expression correlate with a condition. Relief can also be made sensitive to interaction effects (known in genetics as epistasis ). In this episode, Trang Lê joins me to talk about Relief and her version of Relief called STIR (STatistical Infer...

#25 Transposons and repeats with Kaushik Panda and Keith Slotkin 24.09.2018

Kaushik Panda and Keith Slotkin come on the podcast to educate us about repetitive DNA and transposable elements. We talk LINEs, SINEs, LTRs, and even Sleeping Beauty transposons! Kaushik and Keith explain why repeats matter for your whole-genome analysis and answer listeners’ questions. Links: Keith’s paper: The case for not masking away repetitive DNA Questions for this episode on Reddit

#24 Read correction and Bcool with Antoine Limasset 31.08.2018

Antoine Limasset joins me to talk about NGS read correction. Antoine and his colleagues built the read correction tool Bcool based on the de Bruijn graph, and it corrects reads far better than any of the current methods like Bloocoo, Musket, and Lighter. We discuss why and when read correction is needed, how Bcool works, and why it performs better but slower than k-mer spectrum methods. Links: Pre...

#23 RNA design, EteRNA and NEMO with Fernando Portela 27.07.2018

In this episode, I talk to Fernando Portela , a software engineer and amateur scientist who works on RNA design — the problem of composing an RNA sequence that has a specific secondary structure. We talk about how Fernando and others compete and collaborate in designing RNA molecules in the online game EteRNA and about Fernando’s new RNA design algorithm, NEMO , which outperforms all prior publish...

#22 smCounter2: somatic variant calling and UMIs with Chang Xu 29.06.2018

In this episode I’m joined by Chang Xu . Chang is a senior biostatistician at QIAGEN and an author of smCounter2, a low-frequency somatic variant caller. To distinguish rare somatic mutations from sequencing errors, smCounter2 relies on unique molecular identifiers, or UMIs, which help identify multiple reads resulting from the same physical DNA fragment. Chang explains what UMIs are, why they are...

#21 Linear mixed models, GWAS, and lme4qtl with Andrey Ziyatdinov 31.05.2018

Linear mixed models are used to analyze GWAS data and detect QTLs . Andrey Ziyatdinov recently released an R package, lme4qtl , that can be used to formulate and fit these models. In this episode, Andrey and I discuss linear mixed models, genome-wide association studies, and strengths and weaknesses of lme4qtl. Links: Paper: lme4qtl: linear mixed models with flexible covariance structure for genet...

Listen to the the bioinformatics chat podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.