Cohort description and recruitment
A complete of 134 volunteers comprising infants (4–15 months outdated at nursery begin, median 10 months, 18 male, 25 feminine) about to attend the primary yr of nursery college, their dad and mom (29–50 years outdated, median 36 years outdated, 30 male, 39 feminine), siblings (2–21 years outdated, median 2 years outdated, 3 male, 4 feminine) and home pets (n = 5, 2 cats and three canines), and educators (34–56 years outdated, median 38.5 years outdated, 10 feminine) have been recruited and enroled throughout 3 nursery faculties (right here recognized as A, B and C), every with 2 distinct courses, within the municipality of Trento (Italy) in June 2022. The courses inside the similar nursery shared few actions (that’s, child drop-off and pick-up) and areas all through the day, and have been adopted by totally different educators. The protocol of this research was permitted by the Ethics Committee of the College of Trento (protocol quantity 2022-040) and by the Ethics Panel of the European Analysis Council Government Company after analysis of the undertaking (microTOUCH Grant settlement ID 101045015). Upon enrolment, volunteers have been requested to offer knowledgeable consent and full metadata questionnaires. Consent for participation of infants was obtained immediately from dad and mom.
Metadata assortment and group
Date of beginning, intercourse, anthropometric information (weight, top), and antibiotic therapy within the 3 months previous the beginning of the research or supplemented throughout its course, along with info relating to putative contacts with different volunteers previous the start of nursery, have been collected for volunteers of all ages. Metadata particularly collected for infants included gestation size, mode of supply and basic food plan at nursery admission (breast or components milk feeding and weaning date of begin). Grownup individuals have been additionally required to offer info relating to previous or ongoing continual situations and relative therapies, and putative maternal anti-Streptococcus B prophylaxis throughout beginning. Food regimen metadata for infants and adults are detailed within the subsequent part.
Dietary info assortment and evaluation
Briefly, most infants had begun weaning at T01 (weaned n = 38, not weaned n = 2, not out there = 3) and acquired similar strong meals whereas within the nursery. The bulk adopted a combined feeding method throughout weaning, combining strong meals with any sort of milk (combined food plan n = 24, completely strong meals n = 14, not out there = 5). Amongst these infants receiving milk supplementation, feeding varieties have been comparatively balanced (breast-fed n = 9, formula-fed n = 10, receiving each n = 5). Lastly, adults detailed their long-term dietary habits by way of the compilation of the EPIC Meals Frequency Questionnaire (FFQ). FFQs have been used to calculate the wholesome Plant-based Food regimen Index65. High quality and amount of plant-based meals have been derived from FFQs for a complete of 18 meals teams, and divided into quintiles and assigned constructive or unfavorable scores. Members whose consumption exceeded the best quintile acquired a rating of 5, whereas these beneath the bottom quintile acquired a rating of 1. Wholesome plant-based meals acquired constructive scores, whereas much less wholesome or unhealthy plant-based and animal-based meals acquired a unfavorable rating. A last rating was derived by summarizing the scores of every participant. Metadata have been collected and utilized after pseudonymization of volunteers IDs.
Pattern assortment
Pattern assortment started every week earlier than the beginning of the primary time period of nursery (August 2022) and ended after the Christmas holidays (January 2023) for all volunteers. In the course of the first 2 weeks the nursery organized a ‘settling-in part’, through which infants have been step by step launched to the nursery and attended it for about 3 hours per weekday. Within the following weeks, infants attended the nursery for about 8 hours per weekday. All through the time period size (about 14 weeks), stool samples of toddler individuals have been collected weekly (from earlier than nursery admission T01 to on the finish of Christmas holidays T15) by the nursery workers or the researcher within the nursery from nappies saved at room temperature on the identical day of use, utilizing assortment tubes for specimen assortment containing 9 ml of DNA/RNA Defend buffer (Zymo). Pattern assortment was prolonged till the tip of the second time period of the yr (about 30 weeks, ending July 2023) for all donors in group 1 of nursery A, together with infants, dad and mom, educators and pets, sustaining sampling time-point frequencies and modalities. Two follow-up time factors have been collected for all individuals enroled, on the finish of the yr of nursery (July 2023, ‘TA’) and on the finish of the summer season break (August/September 2023, ‘TB’). The samples collected have been moved to the lab and DNA-extracted inside 2 weeks of supply. Samples assortment of infants throughout summer season or winter breaks time factors along with these of siblings and pets have been carried out immediately at residence by the dad and mom and saved at room temperature till the start of nursery (most 2 weeks later). All grownup individuals’ samples have been self-collected following detailed directions, delivered to the lab and processed as beforehand. Educators donated month-to-month, whereas dad and mom collected one further pattern midway the research interval, along with preliminary and last pattern time factors.
DNA extraction and sequencing
After vortex homogenization, DNA was extracted utilizing the DNeasy PowerSoil Professional Equipment (Qiagen), following the instructions of the Human Microbiome Undertaking protocol66. Extra homogenized aliquots have been saved at −20 °C. DNA was quantified utilizing Qubit 2.0 fluorometer (Thermo Fisher Scientific). Sequencing libraries have been ready utilizing the Nextera DNA Library Preparation Equipment (Illumina), as described by the producer’s pointers. The sequencing was carried out on the Illumina NovaSeq 6000 platform following producer’s protocols. The sequencing depth was set at 15 Gbp.
Metagenome high quality management and preprocessing
Stool samples sequences have been pre-processed utilizing the pipeline described at https://github.com/SegataLab/preprocessing. Briefly, metagenomic reads have been quality-controlled and reads of low high quality (high quality rating
Species-level profiling
Profiling on the decision of SGBs was carried out with MetaPhlAn (v4.1)16,68 utilizing the vJun23_202307 markers database and utilizing the –unclassified_estimation parameter (Supplementary Desk 4). SGBs with <0.1% relative abundance in all stool samples have been faraway from taxonomic profiles for calculation of variety indices.
Constructing strain-level phylogenetic bushes
To reliably detect strain-sharing occasions, we augmented our dataset with oral samples from the identical cohort not analysed on this research (n = 342) and extra samples from 16 public longitudinal cohorts. To take action, we queried the curatedMetagenomicData (v3.18)69 for stool samples sequenced at the least at 1-Gbp depth from wholesome westernized human people with at the least 2 time factors per particular person and three such people per dataset. We went by means of the corresponding papers and excluded research involving an intervention between the sampled time factors. Thus now we have included samples satisfying the above standards from the next datasets: ShaoY_2019 (ref. 32), MehtaRS_2018 (ref. 70), VatanenT_2016 (ref. 71), HMP_2019_ibdmdb72,73, BackhedF_2015 (ref. 74), CosteaPI_2017 (ref. 75), YassourM_2018 (ref. 3), KosticAD_2015 (ref. 76), LouisS_2016 (ref. 77), FerrettiP_2018 (ref. 1), HallAB_2017 (ref. 78), WampachL_2018 (ref. 79), NielsenHB_2014 (ref. 80), Heitz-BuschartA_2016 (ref. 81), ChuDM_2017 (ref. 82) and AsnicarF_2017 (ref. 83). In whole, phylogenetic bushes have been constructed utilizing information from 1,405 samples from this cohort and 4,322 samples from further cohorts.
For all of the SGBs detected in our cohort, we queried MetaRefSGB, a microbial genomic database containing >156,000 isolate genomes and >952,000 metagenome-assembled genomes (MAGs) as of vJun23 (ref. 16), for isolate genomes or MAGs from meals sources, and included them within the bushes as references84.
To construct every tree, we included all of the samples for which the SGB was detected by MetaPhlAn. SGBs detected in lower than 20 samples from our cohort have been discarded, leaving 1,363 SGBs. The strain-level phylogenetic bushes have been constructed for every SGB with StrainPhlAn (v4.1)16,68. For 1,107 SGBs, we have been in a position to construct the phylogenetic tree; the remaining 256 SGBs didn’t have a adequate variety of samples with sufficient marker genes with minimal protection as reported by StrainPhlAn.
Evaluation of SSRs
We discarded bushes constructed from alignments shorter than 1,000 nt (n = 111 SGBs discarded). Within the phylogenetic bushes with genomes from a meals supply, we thought-about the ANI on the marker genes (mutation charges in StrainPhlAn) and discarded samples nearer than 99.85% ANI as in all probability coming from meals. When greater than 20% of the samples can be discarded, we dropped the SGB altogether (n = 7 SGBs discarded).
We calculated phylogenetic distances between all samples because the size of the shortest path between the samples alongside the tree branches. We normalized distance distributions inside every tree by dividing by the median distance.
To outline strain-sharing occasions, we calculated a threshold greatest separating the within-individual distribution from the across-individuals distribution of phylogenetic distances. For the within-individual distribution, pairs of samples from the identical particular person sampled most 6 months aside have been thought-about utilizing most one pair per particular person. Among the many doable pairs, we selected the one maximizing the possibilities to sort the pressure of the SGB of curiosity in each samples, and we did so by selecting the pair for which the pattern with the bottom protection for the SGB between the 2 samples was the best amongst all pairs. In case of ties, we maximized the upper estimated protection of the 2. The SGB protection was estimated because the sequencing depth instances the SGB’s relative abundance. For the across-individuals distribution, we choose pairs of samples coming from totally different datasets, one pattern per particular person maximizing the protection. When there have been lower than 50 pairs within the across-individual distribution we discarded the SGB (n = 365 SGBs discarded). When there have been lower than 20 pairs within the within-individual distribution (n = 276 SGBs), we calculated the edge because the third percentile of the across-individual distribution, that’s, setting the anticipated false discovery price to three%. When the within-individual distribution had at the least 20 pairs (n = 309 SGBs), the edge separating the distributions was calculated as maximizing Youden’s index, until the anticipated false discovery price exceeded 5%, through which case we set the edge because the fifth percentile of the across-individual distribution (n = 61 SGBs). After guide curation of the bushes, together with outlier branches elimination, an extra 112 SGBs have been discarded. In whole, we assessed strain-sharing for 512 SGBs.
For every pair of samples, we referred to as a strain-sharing occasion of an SGB when their phylogenetic distance within the corresponding tree was decrease than the corresponding calculated threshold.
For every pattern, we thought-about the SGBs through which corresponding bushes it was positioned, that’s, profiled by StrainPhlAn. Lastly, we outline the SSR because the variety of strain-sharing occasions divided by the variety of SGBs profiled on the pressure degree in each of them by StrainPhlAn. For pairs with lower than 5 SGBs profiled in frequent, we set the SSR as undefined and such pairs have been excluded from SSR evaluation. Pressure retention price is outlined because the within-individual SSR. The within-individual pressure substitute price is outlined because the variety of strains not retained longitudinally amongst retained SGBs over the variety of retained SGBs.
Samples discovered to be contaminated or mislabelled in line with strain-sharing evaluation (and validated with CrocoDeEL (v1.0.6)85) have been eliminated publish hoc (n = 8), lowering the dataset to 1,013 metagenomes. Samples collected throughout antibiotic therapy (n = 26) have been excluded from all following analyses.
Inside-SGB pressure heterogeneity, strain-sharing networks and SGB transmissibility
Inside-SGB pressure heterogeneity (Prolonged Information Fig. 7a) was computed on a per-nursery foundation because the variety of strains of a given SGB current amongst infants at a given time level over the variety of infants having the SGB on the similar time level. A within-SGB pressure heterogeneity of 1 signifies there is no such thing as a strain-sharing amongst infants having the SGB (that’s, all of them have a unique pressure of the SGB).
Pressure-sharing matrices have been used to construct unsupervised strain-sharing networks (Fig. 3b) with the R packages ggraph (v2.2.1) and tidygraph (v1.3.1), the place solely nodes with diploma >0 are proven.
SGB transmissibility (Prolonged Information Fig. 9a) was computed because the variety of particular person pairs sharing the pressure over the full variety of probably callable strain-sharing occasions involving the SGB (that’s, the variety of pairs of people through which the SGB was current and typable in line with StrainPhlAn)8. When the SGB was not current amongst at the least three pairs inside a class, the transmissibility of the SGB inside the class was set as undefined. Differential transmissibility between child–child and child–mom or child–father pairs was assessed for all SGBs having at the least 10 child–child pairs sharing it, with software of a Fisher’s actual take a look at (together with false discovery price management) in case there have been at the least 10 child–mom pairs or 10 child–father pairs sharing the SGB; in any other case the variety of SGB- and strain-sharing occasions have been reported for every group with out software of the take a look at.
Comparability of the contribution of familial and nursery strains with the toddler microbiome
To match the contribution of familial and nursery strains to infants’ microbiome composition, for every child, we computed at every child time level (T01 to T15) the variety of strains shared both with any member of the household or with another toddler of the nursery group, disregarding strains shared with each (until in any other case said). This allowed us to compute the proportion of strains for a given child microbiome that was putatively acquired from both the nursery or household (referred to within the textual content and figures because the ‘proportion of strains acquired’), because the strains completely shared with both one of many two teams of people. Contemplating that relations have been much less densely sampled than infants, we thought-about a most of three samples for every of the opposite infants within the nursery group (baseline T01, halftime T08 and last T15, when out there, emulating the sampling timeline of oldsters), including different infants and household samples to the longitudinal evaluation contemplating the time through which they have been sampled (that’s, on the lookout for strain-sharing solely with samples collected up to now or contemporaneous to the goal time level of the goal child). This in all probability explains the nonlinear will increase of proportion of strains acquired from the nursery group at child T08 and T15 noticed (for instance, Fig. 4c). Furthermore, as a unfavorable management for the significantly bigger variety of people within the nursery group (common n = 7) in contrast with the household (common n = 2), we additionally analysed the strain-sharing dynamics with a random nursery group from one other nursery.
Metagenomic meeting and CRISPR evaluation
MAGs have been generated by means of a beforehand validated metagenomic meeting pipeline17, together with meeting of contigs with MEGAHIT (v1.1.1)86, calculation of contigs protection with Bowtie2 (v2.2.9)67, binning of contigs with MetaBAT2 (v2.12.1)87 and quality-checking of bins with CheckM (v1.1.3)88. Medium- and high-quality MAGs have been recognized following the standards beforehand proposed89 and low-quality MAGs have been discarded. The ANI between MAGs was computed utilizing skani (v.0.2.1)90.
To validate the chain of transmission occasions of the pressure of A. muciniphila SGB9226 proven in Fig. 2a, CRISPR arrays have been recognized from MAGs utilizing MinCED (v0.4.2, default parameters)91. CRISPR spacers and repeats have been extracted from uncooked sequencing reads utilizing Crass (v1.0.1, parameters -d 20 -D 55 -s 20 -S 55 –longDescription)92. Following the identification of 5 CRISPR arrays and 39 CRISPR spacers on this set of MAGs, we regarded for the CRISPR spacers of this pressure within the metagenomic reads of the entire dataset. Though single CRISPR spacers have been discovered within the metagenomic reads of as much as 16/19 samples containing the pressure (Fig. 2a), not one of the 39 CRISPR spacers may ever be discovered within the remaining 627 metagenomes, offering unbiased validation for the trajectory of the pressure within the nursery group.
B. longum strains project to subspecies
To judge the transmissibility of distinct subspecies of B. longum SGB17248 in our dataset, we constructed a StrainPhlAn 4 phylogenetic tree of all 591 B. longum strains recognized in our cohort, which revealed 2 distinct clusters (cluster_1 and cluster_2; Prolonged Information Fig. 10c–e). Cluster_1 consisted completely of strains current in infants and therefore was hypothesized to belong to B. longum subsp. infantis, whereas cluster_2 contained strains from each infants and adults, probably representing different subspecies of B. longum. To definitively assign these strains to subspecies, we succeeded in producing MAGs (as described above) for 262 of the 591 strains after which calculated the ANI between these MAGs and 15 reference genomes representing the three well-known B. longum subspecies93 (5 reference genomes every for subsp. longum, subsp. infantis and subsp. suis). The 24 MAGs from cluster_1 confirmed highest similarity to B. longum subsp. infantis reference genomes (median ANI 98.04%), in contrast with decrease ANI values with subsp. longum (95.60%) and subsp. suis (96.13%). Conversely, the 238 MAGs from cluster_2 confirmed the best genomic similarity to B. longum subsp. longum reference genomes (median ANI 98.87%), with considerably decrease similarity to subsp. suis (96.82%) and subsp. infantis (95.39%). The 591 B. longum strains as profiled with StrainPhlAn 4 have been then assigned in line with their genome-level project.
Blastocystis detection and ST profiling
The presence of Blastocystis in metagenomic samples was assessed utilizing a beforehand validated computational workflow94. Briefly, 9 reference genomes for eight distinct Blastocystis subtypes (that’s, subtype 1 (ST1) (ST1_LXWW01), ST2 (ST2_JZRJ01), ST3 (ST3_JZRK01), ST4 (GCF_000743755 and ST4_BT1_JZRL01), ST6 (ST6_JZRM01), ST8 (ST8_JZRN01) and ST9 (ST9_JZRO01)) have been mapped towards metagenomic reads with Bowtie2 (v2.5). Then, SAMtools (v1.19) and bedtools (v2.30) have been used to compute the breadth of protection of every genome. We reported a pattern to be constructive for a Blastocystis ST if the respective genome had a breadth of protection of at the least 10%.
PCR validation of transmission of an A. muciniphila SGB9226 pressure
To additional take a look at how a lot potential issues with restrict of detection within the metagenomic method may affect pressure transmission inference, we carried out an SGB-specific PCR assay for A. muciniphila (SGB9226) and utilized it to the consultant instance depicted in Fig. 2a. We designed the primers (F-5′-TGACTGGACTCTATTGCCTGAAG-3′ and R-5′-GCCTTTCAATATGCCCTTCGTAC-3′; amplicon size 101 bp) to acknowledge the SGB9226-specific core gene (UniRef90_A0A2N8IRV1; recognized by MetaPhlAn 4 (ref. 16)), utilizing ConsensusPrime (v.1.0, with set consensus similarity 0.8 and consensus threshold 0.95)95 and Primer3 (v. 2.6.1; with set primer measurement 18–28 bp, optimum melting temperature 57–63 °C, GC content material 40–60%, PRIMER_MAX_HAIRPIN_TH 24.0, PRIMER_INTERNAL_MAX_HAIRPIN_TH 24.0, PRIMER_MAX_END_STABILITA 9.0)96. Assay sensitivity was independently evaluated utilizing a spike-in method with DNA from the A. muciniphila sort pressure ATCC BAA-835 (from 1 M right down to 1 single genome copy) into an A. muciniphila-negative faecal take a look at pattern. The assay achieved a restrict of detection equal to a single genome copy of A. muciniphila (Prolonged Information Fig. 6a). PCRs have been carried out utilizing GoTaq G2 Scorching Begin Inexperienced Grasp Combine (Promega) with 500 nM of every primer. The thermal biking programme included an preliminary denaturation at 95 °C for 10 min, adopted by 40 cycles of annealing at 62 °C. Constructive bands have been noticed by electrophoresis on a 2% agarose gel.
Statistical evaluation
Statistical analyses have been carried out in Python (v3.10.12) utilizing libraries scikit-bio (v0.5.9), scipy (v1.10.1) and statsmodels (v0.14.0). Cross-sectional comparisons between teams have been carried out utilizing the Mann–Whitney U-test (two teams) or the Kruskal–Wallis take a look at with publish hoc Dunn checks (a number of teams) for unbiased observations. Cross-sectional dependent observations have been in contrast with a permutation take a look at for medians (or means, when evaluating the variety of shared strains), with P values calculated because the proportion of instances (out of a 1,000 permutations) that the noticed distinction within the median between teams with shuffled labels is equal or extra excessive than the noticed within the appropriately labelled information. Longitudinal comparisons between time factors have been carried out utilizing the Wilcoxon signed-rank take a look at. Jaccard dissimilarity matrices have been computed from taxonomic profiles and compositional variations between teams have been evaluated utilizing permutational multivariate evaluation of variance. When applicable, correction for a number of testing was utilized utilizing the Benjamini–Hochberg process (Padj), with significance outlined as Padj < 0.05.
Reporting abstract
Additional info on analysis design is offered within the Nature Portfolio Reporting Abstract linked to this text.
