This is the kick-off thread for the Planetary Protection AWG @PPawg Bioinformatics sub-group. The first task would be to find a suitable date to meet for most of us.
I would suggest Tuesdays at 6 am PT, 9 am ET, 15:00 CET, always in the week before our monthly PP AWG meeting on Thursday. Next chance would be Tuesday, April 7th. Any time constraints?
As soon as we agree on the date, we should discuss topics and potential projects related to:
Consensus metagenomics pipelines.
Reference genome databases.
AI/ML risk detection algorithms.
Public data/protocol sharing via NASA OSDR and metadata standards, etc.
We are interested in e.g.:
Developing new algorithms
Benchmarking (new or existing)
Integrating and consensus
Applying (New comparisons, space vs cleanroom vs hospital data, large-scale data mining)
bridging with other subgroups (ie. large scale mining of blanks in MGnify in dialog with the nucleic acid sub-group)
I am interested! I’ll be there on behalf of the Space Microbial Culture Collection at Ames, which is currently in the process of ingesting a lot of planetary protection-relevant isolates and developing protocols for sequencing and identification. (So, not metagenomics, but still bioinformatics.) I’d love to be plugged into this AWG and specifically this sub-group. Please include me in invitations and communications. I may not make all meetings but I will try.
Let’s see who’s available to join us tomorrow. Apologies in advance for the last-minute reminder. On the agenda: a round of introductions and a discussion about the first potential projects for our sub-group. Hope to see you there! Cheers, Alex
Hi members of the PP AWG Bioninformatics Sub-Group!
at our next meeting (Tuesday, June 9th 6a PT/9a ET/15:00 CET) we will continue our discussion on the “Low biomass Illumina dataset analysis” from Chelsi Cassilly and Olabiyi Obayomi.
Great to see data in detail @olabiyi! I don’t think it’s a problem that it might take more than an hour to navigate and discuss with this type of thing. Great work.
Hello. It looks like I just missed your most recent meeting. Could you let me know when the next meeting will be and how I can contribute? Thank you and looking forward to working with all of you.
Due to the summer break in July and August, we haven’t yet decided whether a PP AWG Bioinformatics Sub-group Meeting will take place during those months. Our usual schedule for the sub-group is the Tuesday before our PP AWG monthly meeting on Thursday (in the second week of each month). I will provide an update as soon as possible.
Thanks @olabiyi for the discussion about the multicenter study data. I look forward to working with you, Aaron Regberg (@aregberg), Chelsi Cassilly and others to develop the narrative for an upcoming manuscript!
Hi everybody. I apologize for they delay. It took me a little while to find this post again. Here are my thoughts based on what @olabiyi has shared so far.
It seems like there are patterns in this data set that are visible despite evidence for cross contamination between samples. Each center has some number of unique microorganisms present in the data.
The Metaphlan2 classifier is a lot more conservative when identifying reads and is leaving a large portion of the reads, including the spike-in controls in the unidentified bucket.
we should normalize the samples to a uniform read depth and rerun the analysis to confirm the patterns we are seeing in the data.
@nicholas.brereton suggested trying to quantify the amount of cross contamination we are observing between samples. This could be as simple as describing the number of reads assigned to the spike in control measured in samples it was not added to. Is there a more rigorous way to do this?
I think the paper should describe the logistical hurdles associated with collecting samples from multiple centers and analyzing them across multiple research groups. Additionally, it should highlight the importance of using negative controls, mock communities, and spike-in controls to detect sample handling issues and instances of cross contamination. The paper should attempt to quantify the amount of cross contamination observed. The paper should focus on the patterns observed between centers and describe how the use of controls impacts our confidence in this interpretation. Finally, the paper should describe how this data could be used to meet planetary protection requirements and where there are limitations data and tools currently available. I think it makes sense to focus on the output from a single classifier for this effort in order to simplify the data and make the interpretation clear.
I’m interested to hear what everyone else thinks about this approach.
I suggest microbiome folks try the One Codex server. Run samples side by side. We’ve spent 8 years optimizing the outputs to ensure that silly and obviously bad results are dealt with. Outputs such as E coli, Pseudomonas aeriginosa, Bradyrhizobium, which are a mess in NCBI have been corrected. Obviously for most simple these organisms are not present. Unless of course you’re doing wastewater… This is not the case for Kraken2 and MetaP. There’s a lot of reasons for this that I’m not going to go through here but give it a try. Obviously it is not for 16S since that is a total weight of time anyway.
Scott Tighe
Senior Research Associate | University of Vermont | 802-999-6666
Environmental Microbiome Engineering Research Group (EMERG Laboratory)
Lab -Votey Rm 369 |Office- Votey 213F
33 Colchester Ave | Burlington, VT 05405
Soil Health Research and Extension Center (SHREC Laboratory)
Jeffords Building (Lab)
63 Carrigan Dr | Burlington, VT 05405