Bacterial genomic datasets compiled for this research
A number of bacterial genomic datasets have been compiled for this research, and totally different choices of those datasets have been used, as acceptable, for the analyses described herein. We refer to those datasets as: ‘NCBI full genomes’, ‘NCBI long-read assemblies’, ‘NCBI short-read E. faecium assemblies’, ‘SHC long-read Enterococcus isolates’ and ‘SHC long-read metagenomes’. We additionally outline the ‘high-contiguity Enterococcus genomes’, which include Enterococcus genomes from (1) NCBI full genomes, (2) NCBI long-read assemblies, and (3) SHC long-read Enterococcus isolates. Datasets have been mixed with out deduplication.
To assist readers in decoding our manuscript and for transparency, Supplementary Desk 1 accommodates (1) an in depth description of bacterial genomic datasets compiled for this research, (2) descriptions of any further filtering for particular analyses, and (3) an outline of how these datasets are utilized in every principal determine and Prolonged Knowledge determine.
NCBI long-read assemblies: long-read dataset curation, preprocessing, meeting and high quality management
Candidate isolate long-read sequencing datasets have been recognized on NCBI with the next search standards on 29 February 2024: ((((Enterococcus[Organism]) OR (Staphylococcus[Organism] OR (Klebsiella [Organism]) OR (Acinetobacter [Organism]) OR (Pseudomonas [Organism]) OR (Enterobacter [Organism]) OR (Escherichia [Organism]) OR (Salmonella[Organism]) OR (Streptococcus[Organism])) AND ‘oxford nanopore’[Platform])). Datasets for Shigella and Bordetella have been queried individually on 28 October 2024 in the identical approach. Metadata for these accessions was downloaded with the Sequence Learn Archive (SRA) Run Selector interface. Datasets with lower than 50 Mb of complete sequencing or lower than 2 kb common learn size weren’t included for additional evaluation. Uncooked reads have been downloaded utilizing SRA instruments prefetch and fasterq-dump, and assembled with flye56 (v.2.9) utilizing the –nano-raw methodology. Genome taxonomy was verified with GTDB-Tk57 (v.2.4) with database launch v.220.
Assemblies have been then filtered for acceptable contiguity (≤20 contigs), genome dimension (±2 s.d. per genus), protection (≥40×) and with recognized species-level taxonomy. We excluded the GTDB-Tk taxonomy end result for the Shigella datasets as a result of the GTDB-Tk database doesn’t embody Shigella genomes and thus can not classify them past E. coli. Lastly, we filtered the set of genomes to incorporate solely species with greater than ten samples on this and the NCBI full genomes (see ‘NCBI full genomes’ under) mixed. The ultimate set of genomes with meeting and taxonomy info is offered in Supplementary Desk 2. We seek advice from this dataset within the manuscript as NCBI long-read assemblies. Related supply code and information will be present in our repository (‘Code Availability’).
NCBI full genomes
The NCBI command line instrument ncbi-datasets-cli (v.16.22.1)58 was used to obtain all obtainable NCBI genomes with the ‘full’ meeting stage from the next taxa on 3 July 2024: Enterobacter, Staphylococcus, Klebsiella, Acinetobacter, Pseudomonas, Enterococcus, Escherichia, Salmonella, Streptococcus, Shigella and Bordetella. For instance, the next command was used to obtain Enterococcus assemblies: datasets obtain genome taxon enterococcus –assembly-level full –include genome,gff3 –dehydrated –filename enterococcus.zip. Be aware, this command collects a reference to the related genomic information. Knowledge have to be downloaded utilizing the ‘rehydrate’ command. After rehydrating, duplicate assemblies have been eliminated, preferring the RefSeq over GenBank because the supply database. Assemblies have been then filtered based on the identical contiguity, dimension, protection and pattern depend thresholds as above, and the ultimate set is likewise obtainable in Supplementary Desk 2. We seek advice from this dataset within the manuscript as NCBI full genomes. Related supply code and information will be present in our repository (‘Code Availability’).
SHC long-read Enterococcus isolates: assortment and sequencing
A complete of 279 bloodstream an infection isolates recognized by biotyping with a MALDI-TOF primarily based Bruker Biotyper as E. faecalis (n = 170) or E. faecium (n = 109) by the Stanford Well being Care Scientific Microbiology Laboratory between the durations 1 October 2020 to 31 December 2021 and 1 January 2023 to 31 December 2023 have been collected underneath a protocol accepted by the Stanford College Institutional Overview Board (Protocol no. 64591; Principal Investigator: A. Bhatt, Committee IRB-61). Knowledgeable consent was obtained from all human analysis individuals. An extra eight isolates collected in 2024 recognized as atypical Enterococcus as above have been saved as glycerol shares. Metadata, together with date, website of assortment and related affected person info, have been extracted. Every isolate was cultured aerobically for about 16 h with shaking at 37 °C in 3 ml Mind Coronary heart Infusion (BHI) broth (Sigma-Aldrich). Then, a 1:100 subculture into 10 ml BHI broth was ready. Subcultures have been allowed to achieve an OD of roughly 1.0. The subculture was centrifuged at 3,000g for 10 min at room temperature, resuspended in PBS, then centrifuged once more. The PBS was poured off and the next pellet was resuspended in 1 ml 1× Zymo DNA/RNA Protect (Zymo, catalogue no. R1100-250) and shipped to Plasmidsaurus for Oxford Nanopore Applied sciences (ONT) ‘Normal Bacterial Genome with DNA extraction’ sequencing. After inspection of the assemblies, one E. faecium isolate (EF_B_40) was mischaracterized as E. faecalis and one E. faecalis isolate (FM_B_40) was mischaracterized as E. faecium, on the premise of biotyping. Of the 170 E. faecalis isolates, 167 have been confirmed as E. faecalis. Of the 109 E. faecium isolates, 107 have been confirmed as E. faecium with commonplace high quality management carried out by Plasmidsaurus. Lastly, assemblies have been filtered based on the identical contiguity and protection thresholds as above, leading to 166 E. faecalis and 100 E. faecium genomes, respectively. We seek advice from this dataset within the manuscript as SHC long-read Enterococcus isolates.
IS counting
ISEScan59 (v1.7.2.3) was run with default parameters. Every ensuing {pattern}.csv file was used to depend IS parts per pattern, per contig and per IS household. Abstract outcomes from NCBI long-read assemblies and NCBI full genomes are included in Supplementary Desk 2, and abstract outcomes from SHC long-read Enterococcus isolates (E. faecium and atypical Enterococcus) are included in Supplementary Desk 3.
Meeting annotation
Genomes have been annotated with Bakta60 (v.1.9.1–v.1.10.3) utilizing database v.5.1 with the next flags –skip-tmrna –skip-ncrna –skip-ncrna-region –skip-crispr –skip-sorf –skip-gap –keep-contig-headers. Output recordsdata, together with .ffn, .faa and .gff3, have been used for subsequent analyses.
Willpower of IS flanking areas and limits
To research the sequence options of high-copy IS parts in E. faecium we used customized scripts to establish the IS ends, yielding transposase flanking and boundary areas. From the evaluation described within the previous sections, we recognized 17 IS transposases with at the least one copy per genome on common. Of those, we exclude the three IS3 household sequences as they’re annotated as two coding sequences (CDSs) (orfA/B). The ensuing 14 sequences comprise 30,847 copies throughout the 593 E. faecium genomes surveyed. In short, we extracted 500 bp of flanking sequence on both aspect of the transposase CDS. For a given transposase variant, as their CDS is an identical, we count on the IS flanks to be extremely conserved. Nevertheless, as we’re contemplating variants from many comparable genomes, we should account for identification by descent the place sequence context past the true IS boundary is repeated. Subsequently, we first deduplicate the flanking sequence. Then, per transposase variant, a a number of sequence alignment was generated for the set of deduplicated flanks with MAFFT61 (v.7.526). After collapsing positions whose consensus is ‘GAP’ (representing low-frequency insertions), per base entropy was calculated as bits (log2-scale) with scores approaching 0 representing extremely conserved websites. The flank is set because the (longest) area of steady low entropy (<0.5), exempting quick (≤3 bp) high-entropy segments representing SNVs or quick indels. Sequence info for these IS, together with their consensus flank lengths, is displayed in Prolonged Knowledge Fig. 4c. Related supply code will be present in our repository (‘Code Availability’).
A number of sequence alignment of IS transposase ORFs
Transposase protein sequences from the most typical IS parts in E. faecium (Fig. 1c) have been aligned utilizing MAFFT (v.7.526) with default parameters. This excludes IS3 parts as they aren’t encoded as single ORFs. A maximum-likelihood phylogeny was inferred with IQ-TREE2 (ref. 62) (v.2.2.0) underneath the VT+F+G4 substitution mannequin, chosen based on the Bayesian info criterion, with 1,000 ultrafast bootstrap replicates. Department lengths characterize the anticipated variety of amino acid substitutions per website. The ensuing tree was visualized with the R bundle ggtree (v.3.10.1).
Frequency and distribution of IS transposase variants in E. faecium
For the identification of distinctive IS transposase ORF variants and quantification of their frequencies in E. faecium, we analysed the 593 E. faecium genomes from NCBI full genomes and SHC Enterococcus isolates. We excluded the NCBI long-read assemblies from this evaluation as a result of solely 5 out of 174 datasets (Prolonged Knowledge Fig. 1) have been sequenced with the more moderen R10.4.1 nanopore chemistry. Earlier-generation ONT information are identified to exhibit increased charges of homopolymer errors63, which don’t notably hamper hidden-Markov-model-based approaches akin to ISEScan however would complicate correct counting of actual copies of transposase nucleotide sequences.
Genome annotations have been carried out utilizing Bakta, and the ensuing GFF recordsdata have been used to extract all CDS options longer than 500 nucleotides and labelled as ‘transposase’. These have been then consolidated right into a non-redundant transposase database, enabling us to quantify the frequency of every distinctive transposase nucleotide sequence throughout the dataset. We recognized the six most considerable IS transposase households—ISL3, IS30, IS256, IS6, IS3 and IS110—and visualized their distribution (Prolonged Knowledge Fig. 4d). Related supply code will be present in our repository (‘Code Availability’).
geNomad prediction of plasmid contigs
To foretell whether or not a contig might be from a chromosome or plasmid, we used geNomad64 (v.1.8.0, v.1.6 DB) with the end-to-end workflow. Contigs predicted as chromosomal with a rating of >0.6 have been included in analyses as chromosomes. Contigs predicted as a plasmid with a rating of >0.7 have been included in analyses as plasmid. Contigs not reaching these thresholds have been thought-about undetermined and excluded from analyses the place plasmid predictions have been essential (Fig. 1c and Prolonged Knowledge Fig. 3). Details about the variety of contigs reaching these thresholds for every genome from the NCBI full genomes and NCBI long-read assemblies are included in Supplementary Desk 2 and are visualized in Prolonged Knowledge Fig. 3i.
IS counts amongst ESKAPEE lineages
To manage for the chance that sure lineages drive species-level ends in Fig. 1b, IS counts have been measured amongst ESKAPEE taxa lineages. All pairwise genomic distances have been calculated inside every ESKAPEE taxon utilizing Mash65 (v.2.3) with 100,000 (-s), size 21 okay-mer (-k) sketches. Genomes have been then grouped into clusters outlined by <0.01 pairwise distance throughout all members (Prolonged Knowledge Fig. 2).
IS growth in enterococci
Genomes from the high-contiguity Enterococcus assortment and the 32 reference genomes described by Lebreton and colleagues29 (2,134 in complete) have been categorised phylogenetically with de_novo_wf from GTDB-Tk v.2.4.0, Launch 220. Following Lebreton and colleagues29, the time period g__Vagococcus was used to outline the outgroup. A desk summarizing the IS counts as outlined by ISEScan for every genome, in addition to the Newick and tree.dist recordsdata output by de_novo_wf have been processed with customized scripts. Along with 4 genomes being excluded by GTDB-Tk, genomes have been assigned to a consultant by the minimal phylogenetic distance. A complete of 18 genomes have been excluded for having a minimal distance >0.01, yielding 2,112 assigned genomes. IS counts per household have been summarized for every set of genomes assigned to a given consultant. In complete, 785 genomes have been assigned to the representatives EnGen0015, Aus0004 and EnGen0043, outlined as E. lactis, clade A1 (healthcare-associated) and clade A2, respectively. To visualise counts per IS household throughout enterococci (Fig. 2b), the imply depend of every IS household for the 32 units of genomes was calculated and a file containing these values was uploaded to iToL66. For higher visualization of Enterococcus, Lactococcus garvieae was excluded from the tree. Related supply code and information will be present in our repository (‘Code Availability’).
ISL3 counts throughout E. faecium lineages
To research whether or not sure lineages of E. faecium have been enriched for ISL3 copy, ISL3 copy was related to a phylogenetic tree of all E. faecium genomes from the high-contiguity Enterococcus assortment. In short, a pangenome was generated with panaroo67 (v.1.5.2, –core_threshold 0.99). The ensuing core genome alignment was enter into iqtree2 (v.2.2.0) to construct a phylogenetic tree utilizing a GTR+G substitution mannequin and 1,000 bootstraps. The utmost-likelihood tree was visualized in iToL. The tree was midpoint rooted between E. lactis and clades A1 and A2. For every of the 785 genomes, ISL3 copy quantity as measured beforehand by ISEScan was uploaded to iToL. Moreover, we assigned STs with multi-locus sequence typing (MLST)68 (v.2.23.0; obtainable at https://github.com/tseemann/mlst) to every genome utilizing the scheme efaecium. Clade info was uploaded to iToL. Along with E. lactis and clade A2 E. faecium, widespread STs (17, 18, 80, 117, 203 and 1,421) have been additionally visualized.
Correlation of antibiotic resistance and IS copy in E. faecium
To foretell antibiotic resistance genes (ARGs), AMRFinder69 (v.4.0.23, DB 2025-07-16.1) was run on 785 E. faecium genomes from the high-contiguity Enterococcus assortment with further parameters -plus and -organism ‘Enterococcus faecium’. Hits have been deduplicated by choosing the very best identification per ORF coordinate. Then, distinctive ARGs per genome have been counted. ARGs have been then associated to IS counts per genome; 65 genomes predicted to be E. lactis as outlined above in ‘IS growth in enterococci’ have been thought-about a single ‘Non-clade A1/A2’ class. The remaining 720 predicted clade A1/A2 genomes have been transformed right into a categorical variable on the premise of IS copy, for every of the six most typical IS households as in Prolonged Knowledge Fig. 4d. The genomes have been separated into quintiles, with the bottom quintile as Q1 (fewest IS) and the very best quintile as Q5 (most IS).
Defining human-infectious NFF Enterococcus
NFF Enterococcus inflicting human an infection have been primarily based on the Australian Enterococcus Surveillance Consequence Program30. In the course of the 10 years from 2013 to 2022, 11,681 Enterococcal bacteremia isolates have been collected and recognized. Any Enterococcus species answerable for at the least ten bacteremia isolates (common of 1 per 12 months) in that assortment is taken into account right here as human-infectious. These species are E. faecalis (6,319), E. faecium (4,712), E. casseliflavus (198), E. gallinarum (166), E. avium (107), E. raffinosus (60), E. hirae (45), E. lactis (29) and E. durans (27).
NCBI short-read E. faecium assemblies: curation of isolate datasets for ISL3 growth timeline
A listing of accessions was obtained by querying the NCBI e-utilities utility programming interface, leading to a set of 21,376 Illumina-sequenced E. faecium genomic whole-genome sequencing information with obtainable assortment 12 months. This set of accessions was filtered to paired-end information with at the least round 20× estimated protection after which downsampled to ≤500 samples per 12 months, leading to a remaining set of 6,747 samples. Subsequent, we ran customized Nextflow70 pipelines to course of the samples. In short, uncooked reads have been downloaded, deduplicated, trimmed, coverage-filtered/downsampled (20–150×) and assembled, after which FastANI71 (versus clade A1 [Aus0004], clade A2 [EnGen0015] and E. lactis [EnGen0043] references) and CheckM2 (ref. 72) have been run on every meeting. Lastly, assemblies have been filtered: (1) FastANI closest reference = A1, min_dist ≤ 0.015 and ≥97% common nucleotide identification (ANI); (2) genome dimension inside µ ± 2 s.d.; (3) CheckM2 ≥ 99% completeness and ≤5% contamination; (4) contig N50 ≥ 25 kb. This resulted in 5,422 remaining short-read assemblies; we seek advice from this dataset within the manuscript as NCBI short-read E. faecium assemblies. Related supply code and information will be present in our repository (‘Code Availability’).
NCBI short-read S. epidermidis assemblies: curation of isolate datasets for IS growth timeline
As with E. faecium, a listing of accessions was obtained querying the NCBI e-utilities utility programming interface, ensuing within the set of 5,938 Illumina-sequenced S. epidermidis genomic whole-genome sequencing information with obtainable assortment 12 months. To counterpoint for hospital-adapted isolates and keep away from the potential confounding affect of pores and skin commensals, we excluded isolates with sampling metadata indicating superficial physique websites (for instance ‘pores and skin’, ‘scalp’, ‘hand’, ‘ear’). We additionally required ≥99% ANI to the S. epidermidis RefSeq reference (GCF_006094375.1) and utilized further high quality management thresholds analogous to these in our E. faecium evaluation. This resulted in 1,604 remaining short-read assemblies after making use of all filtering steps. Related supply code and information will be present in our repository (‘Code Availability’).
Protection-based IS depend estimates in short-read information
To quantify IS abundance in these short-read assemblies, we first annotated IS current on particular person contigs with ISEScan. Contigs shorter than (600 × variety of predicted IS on the contig) weren’t counted to make sure contigs which might be putatively solely a part of an IS should not counted. This limits the inclusion of noisy protection information from very quick contigs, in addition to the potential for double counting fragmented IS parts. This most likely additionally results in conservative estimates of IS depend, which we want to the choice. For every meeting, chromosomal protection was approximated because the length-weighted median protection of contigs exceeding 5 kb in size. We then calculated a relative protection for every contig by dividing the contig protection by the chromosomal protection. The variety of IS copies on every contig was multiplied by its estimated protection, and these values have been summed subsequently throughout all included contigs to acquire the ultimate per household IS depend per meeting.
To make sure our IS counting heuristic from fragmented assemblies was acceptable, particularly to keep away from differential bias in direction of sure IS households, we benchmarked our strategy towards extremely contiguous genome IS counts. We generated paired short- and long-read information for 117 samples (37 E. faecium; 80 E. faecalis) from our SHC assortment and in contrast IS counts per household and per species (Fig. 3b and Prolonged Knowledge Fig. 5). These outcomes present that our IS counting heuristic within reason correct throughout IS households and that this isn’t particular to E. faecium. The outcomes additionally spotlight the next caveats of the heuristic: (1) it tends in direction of being barely conservative (due partially to ignoring quick noisy contigs); (2) it’s much less exact, though not much less correct, for higher-copy IS households and (3) it doesn’t carry out as nicely for IS6, which tends to be plasmid-borne (Fig. 3b). This end result is no surprise due to the larger issue in resolving repetitive/plasmid sequences in short-read information. Consequently, we don’t embody analyses of IS6 parts from short-read-based assemblies.
To manage for the chance altering IS abundance over time may very well be defined by latest isolates being disproportionately medical in origin, whereas older isolates may derive from nonclinical sources, we evaluated a number of IS households. Healthcare-associated E. faecium strains have been reported beforehand to be enriched for IS256, IS30 and IS3 household parts. Consequently, if this sort of sampling bias have been driving our findings, we might count on all IS households related to medical strains to comply with the identical temporal sample (that’s, low in older and excessive in more moderen samples).
Related supply code and information will be present in our repository (‘Code Availability’).
Trendline of IS copy quantity over time
To discern whether or not there’s a normal pattern in altering IS copy quantity over time (uncooked information proven in Prolonged Knowledge Fig. 5), whereas limiting short-term fluctuations, we computed the typical genomic IS copy quantity per 12 months over a 5-year sliding window (95% confidence interval, imply ± t × s.e.), and plotted a trendline utilizing domestically weighted scatterplot smoothing (LOWESS) (smoothing fraction 0.25). Related supply code will be present in our repository (‘Code Availability’).
PopPUNK clustering of E. faecalis and E. faecium genomes for PanGraph enter
To characterize structural variation amongst a set of almost clonal genomes, we returned to the SHC long-read Enterococcus isolates as described above. With PopPUNK73 (v.2.6.7), we clustered 722 E. faecium genomes from the high-contiguity Enterococcus assortment with the v.2 full E. faecium PopPUNK database downloaded from https://ftp.ebi.ac.uk/pub/databases/pp_dbs/Enterococcus_faecium_v2_full.tar.bz2. These genomes are the 720 clade A genomes as in Fig. 2a, with FM003 and FM057 included. They have been excluded beforehand for protection decrease than our 40× threshold (39× and 31×). We selected to incorporate them right here as a result of they’re each single-contig chromosomal assemblies. Clustering was performed with the poppunk_assign command with the extra flag of run-qc. PopPUNK’s poppunk_visualize command was used to generate a Newick file, viz_core_NJ.nwk, describing the ensuing phylogeny. The output tree was visualized with iToL. We assigned sequence sorts with MLST (v.2.23.0) to every genome utilizing the scheme efaecium. The SHC hospital-endemic lineage corresponds to ST117 and fashioned a single cluster. The 66 genomes within the cluster, all from the SHC assortment, have been chosen for additional evaluation. Equally, the 1,083 E. faecalis genomes from NCBI full and SHC long-read Enterococcus isolates have been assigned to the v.2 full E. faecalis PopPUNK database downloaded from https://ftp.ebi.ac.uk/pub/databases/pp_dbs/Enterococcus_faecalis_v2_full.tar.bz2. None have been eliminated by high quality management. We tried to assign STs as above utilizing the scheme efaecalis. Right here the biggest cluster akin to ST179 contained 57 genomes, 44 of which have been SHC isolates. The 44 genomes from the SHC assortment have been chosen for additional evaluation.
Filtering genomes and stitching contigs for enter into PanGraph
For every of the SHC genomes in these clusters, we extracted the contigs predicted to be from the chromosome by geNomad as described earlier than. As PanGraph74 requires single-contig inputs, we filtered for genomes with 5 or fewer contigs, yielding 66 E. faecium and 44 E. faecalis genomes. Multi-contig genomes have been stitched collectively utilizing 2 kb ‘N’ spacers within the order output by Flye.
PanGraph evaluation of structural variation within the SHC E. faecium ST117 and E. faecalis ST179 clusters
For every of the 2 clusters outlined above, we tailored the ‘Structural genome evolution’ workflow as outlined by Molari and colleagues39 and located at https://github.com/mmolari/structural-evo.
In short, the genomes have been processed with PanGraph (v.0.7.1) to construct a pangenome graph, which defines all ‘core blocks’ (areas of shared homology throughout all genomes). Areas of structural variation are captured as edges between consecutive core blocks, and the set of sequences that may happen between two core blocks is outlined as a ‘junction’. PanGraph was rerun on every junction (with its flanking core blocks), and the ensuing ‘subgraph’ defines every distinctive association of homologous blocks by the junction as a novel ‘junction path’. Thus, for a given junction, the set of distinctive junction paths neatly describes the extent of structural sequence variation occurring within the cluster for a given structurally variant locus.
Subsequent, the set of junction subgraphs was analysed with a customized Python workflow: first, we confirmed that the spacers (2 kb ‘N’ as outlined above) all the time occurred totally within the accent area of junctions. This was anticipated, as disruptions in meeting contiguity are most likely attributable to structural variation inside the pattern inhabitants itself. We eliminated any spacer-containing path from downstream evaluation as their true sequence size and genotype is unclear. Subsequent, for every junction, the workflow enumerates the distinct paths inside the subgraph and characterizes the accent genome of every non-empty path, utilizing the Bakta-derived annotations (described above) for every supply genome (Supplementary Desk 4). In ST117, the most typical sort of junction is a ‘binary’ IS insertion. In that case, the junction contains two paths, by which one path is empty (that’s, there is no such thing as a accent area between the core blocks), and the opposite path accommodates solely an IS component transposase annotation. IS-only junctions containing greater than two paths and/or a number of IS households additionally occurred. Fewer giant and extra complicated junctions occurred as nicely, together with giant indels and phages, though we didn’t characterize these intimately. Related supply code and information will be present in our repository (‘Code Availability’).
Cell genetic parts have been reported beforehand as necessary to E. faecalis virulence. The evaluation of E. faecalis ST179 served as a management for charges of structural variation over the identical interval in a associated organism with comparable hospital epidemiology and taxonomic relatedness. The E. faecalis samples have been collected from the identical surroundings, location and time-frame, and have been processed and analysed in an an identical method (Supplementary Desk 4).
For compatibility with our high-performance computing cluster, the PanGraph workflow was tailored to work for Snakemake (v.8.18.2). The workflow was run per cluster with present installations of ISEScan (v.1.7.2.3), DefenseFinder75 (v.1.3.0), geNomad (v.1.8.0) and IntegronFinder76 (v.2.0.5).
Technology of circos plots of IS variants in PanGraph clusters
The isolate with earliest assortment date, EF006 for the ST179 E. faecalis cluster and FM001 for the ST117 E. faecium cluster, was chosen because the reference coordinate system. The set of IS-only junctions was filtered to ‘backbone-only’ junctions. In different phrases, IS insertions occurring in junctions not shared by all genomes (for instance, inside different junctions or between core blocks which might be inverted for some genomes) weren’t proven, as their place on the core spine is ambiguous. Every junction was given a hard and fast width of three kb, as smaller sizes weren’t seen on the plot. Related supply code and information will be present in our repository (‘Code Availability’).
ISL3 gene disruption and intergenic distance evaluation in ST117
The E. faecium ST117 PanGraph evaluation outlined 100 ISL3 junctions (that’s, IS-only junctions containing solely ISL3 household transposase annotation(s)) (Fig. 4c). Inspection of those junctions revealed that the overwhelming majority are certainly parsimoniously defined by insertions of a number of ISL3 parts within the area. To establish candidates for intra-genic disruptions, we thought-about junctions the place an annotated gene within the ‘empty’ path (missing the ISL3 insertion) overlapped totally or partially with the accent area of a ‘stuffed’ path. A junction was counted as gene-disruptive provided that it contained such an overlapping gene and if that gene’s annotation was additionally lacking or truncated within the stuffed path (Fig. 4h). The remaining ISL3 junctions have been presumed to be intergenic. Of those, we thought-about each stuffed (ISL3-containing) path for each junction and extracted the gap from the three′ finish of the transposase to the following nearest gene, in addition to its relative orientation. Paths that contained a number of accent ISL3 transposases have been counted provided that they have been all oriented the identical approach, in any other case the orientation was thought-about ambiguous and the trail was excluded from evaluation (Fig. 4i). As a management, the identical analyses have been performed on the 51 IS30, IS256, IS110, IS3 and IS6 junctions as a mixed set (‘Different’ in Fig. 4h,i).
E. faecium domination longitudinal pattern curation
To establish a set of metagenomic samples that have been (1) clinically related, (2) longitudinal and (3) technically appropriate for assembling full, closed E. faecium genomes immediately, we chosen longitudinal circumstances from our laboratory’s biobank of stool samples from folks present process HCT with excessive relative abundance of E. faecium. As beforehand described44,45, these stool samples have been collected underneath protocols accepted by the Stanford College Institutional Overview Board (Protocol nos. 8903; principal investigator: D. Miklos, co-investigator: A. Bhatt and T. Andermann, Committee IRB-3 and 42053; principal investigator: A. Bhatt, Committee IRB-8). Knowledgeable consent was obtained from all human analysis individuals. On the premise of earlier metagenomic sequencing outcomes, samples with at the least 10% relative abundance of E. faecium have been thought-about. Any sufferers with at the least two such samples have been then recognized and people samples ‘dominated’ (outlined as ≥10% relative abundance) by E. faecium have been chosen for long-read metagenomic sequencing, yielding 34 samples from 14 sufferers (vary, two to 5 samples per affected person).
Metagenomic excessive molecular weight DNA extraction
To extract excessive molecular weight DNA from stool samples in preparation for metagenomic long-read sequencing, we used the QIAamp PowerFecal Professional DNA Package (Qiagen, catalogue no. 51804). In short, working in a biosafety hood, from every stool pattern stored on dry ice, roughly 200 mg stool (two to a few biopsy punches) have been added to the bead beating tube. Bead beating tubes have been added in units of 12 to a Vortex-Genie 2 with the horizontal Vortex Adapter described within the extraction equipment producer’s directions and vortexed at most velocity for 12 min at room temperature. All subsequent steps have been in step with producer directions utilizing centrifugation. The ultimate DNA was eluted in 100 µl C6 buffer and incubated at room temperature for 15 min earlier than centrifugation. DNA focus was measured by NanoDrop. DNA was then cleaned up with the Genomic DNA Clear & Concentrator-10 equipment (Zymo Analysis, catalogue no. D4011) based on producer directions for big genomic DNA. In short, 200 µl of DNA binding buffer was added to 100 µl of DNA eluate from earlier than. DNA was eluted in 25 µl nuclease-free water. After clean-up, by Nanodrop, DNA concentrations ranged from 7.5 ng µl−1 to 768.7 ng µl−1 per pattern. Two samples had each DNA concentrations under 25 ng µl−1 and purity values (OD260/OD230) lower than 1.5 and weren’t sequenced. DNA high quality was confirmed with Genomic DNA ScreenTape Evaluation (Agilent, catalogue nos. 5067-5365 and 5067-5366) on the Agilent 4150 TapeStation system. Samples had peak concentrations at DNA lengths larger than 20 kb. The lack of one pattern, p11580_03 led to a singleton, however we continued with sequencing of p11580_06. Subsequently 32 samples from 14 sufferers (vary 1 to 4) have been sequenced as described under.
SHC long-read metagenomes: ONT sequencing, basecalling, high quality management and meeting
Two pooled libraries containing 16 samples every have been ready with the Oxford Nanopore Native Barcoding 96-kit (catalogue no. SQK-NBD114.96) based on producer’s directions. In short, 400 ng of enter DNA per pattern was ready, diluting the pattern in nuclease-free water if essential. DNA focus was confirmed for every pattern with the Qubit DNA Excessive Sensitivity equipment at every checkpoint. On the premise of earlier use of this strategy for metagenomic sequencing of stool samples in our laboratory, and on TapeStation outcomes, we anticipated our fragment size to be between 7 kb and 10 kb and due to this fact loaded 50 fmol of every library. Libraries have been loaded onto separate R10.4.1 PromethION move cells (FLO-PRO114M) based on the producer’s directions and sequenced concurrently on a PromethION 2 solo instrument. We anticipated assembling full E. faecium genomes with at the least 40× protection. Subsequently, we estimated that to achieve 40× protection of a 3-Mb E. faecium genome making up 10% relative abundance of the pattern, we would wish 1.2 Gb of sequencing per pattern. To extend the chance of reaching this threshold for every pattern and accounting for library imbalance, we let sequencing proceed till each libraries reached at the least 2.5 Gb per pattern of complete sequencing. We obtained 60 Gb and 40 Gb of sequencing for every library, respectively.
Uncooked.pod5 recordsdata have been transferred to our computing cluster and basecalling was carried out with Oxford Nanopore’s open-source Dorado basecaller (v.0.7.3) obtainable at https://github.com/nanoporetech/dorado. The next command was used: dorado basecaller —kit-name SQK-NBD114-96 -b 1024 —trim ‘all’ sup@v5.0.0 {pod5} > {base known as}.bam. Output bam recordsdata have been demultiplexed and emitted as.fastq recordsdata with: dorado basecaller —kit-name SQK-NBD114-96 -b 1024 —trim ‘demux’. Of 32 pattern barcodes, 31 had at the least 1 Gb of sequencing after demultiplexing. One barcode akin to p11569_01 had lower than 10 Mb of sequencing and was not thought-about additional. Sadly, this made p11569_02 a singleton. Finally, 31 samples from 14 sufferers (vary 1 to 4) have been included for metagenomic meeting (Supplementary Desk 5).
NanoPlot77 (v.1.41.6) was used to guage sequencing high quality per pattern with default parameters. All 31 samples had median learn high quality of at the least Q22.1 (22.1–24.2). Two samples, p11463_06 and p11580_06 had 9,150 and 11,220 reads, respectively. Different samples had at the least 56,643 reads (56,643–659,377). The minimal learn N50 for these 29 samples was 7.77 kb. Metagenomic meeting was carried out with Flye78 (v.2.9.2) on all 31 sequenced samples utilizing the next command: flye –nano-hq {reads} —meta —genome-size 200 m —keep-haplotypes —no-alt-contigs (Supplementary Desk 5). We seek advice from this dataset within the manuscript because the SHC long-read metagenomes.
Taxonomic profiling of assembled contigs with sourmash
To establish E. faecium contigs from the metagenomic assemblies, all contigs 5 kb or longer have been profiled taxonomically with sourmash79 (v.4.8.11) utilizing the successive instructions: sourmash sketch dna -p okay = 51,abund –singleton {sample1}.fa and sourmash scripts fastmultigather -k 51 -m DNA –threshold-bp 5000 {samples}.zip gtdb-rs214-reps.k51.zip. All contigs predicted as g_Enterococcus;s_faecium longer than 2.5 Mb, and with at the least 80× protection have been thought-about ‘full’ chromosomal contigs. One such contig was recognized for 28 of 31 (p11463_06, p11569_02, p11580_06 didn’t) efficiently sequenced samples as described above, which resulted in 12 longitudinal circumstances (Prolonged Knowledge Fig. 8).
Cut up k-mer evaluation of E. faecium assemblies
To find out circumstances of ‘isogenic’ E. faecium chromosomes such that inside pressure heterogeneity and structural variation may very well be assessed clearly, the relatedness of the 28 longitudinal E. faecium chromosomal contigs was thought-about with SKA80 (v.1.0). All pairwise distances and clusters of extremely associated contigs have been recognized utilizing the next instructions: first, ska alleles -f inputs.txt after which, ska distance -f skfs.txt -o outcomes -S -s 40, the place inputs.txt contained paths to the 28 E. faecium chromosomal contig fasta recordsdata. Sufferers with longitudinal samples with fewer than ten SNPs per megabase or increased than 99.999% identification over time have been thought-about for longitudinal within-patient analyses. This threshold excluded p11342.
Inside-sample and within-patient structural variant calling
To evaluate within-sample E. faecium inhabitants heterogeneity, reads from every pattern have been aligned to their very own metagenome meeting utilizing ngmlr81 (v.0.2.7) and default ONT sequencing presets. Structural variants have been known as with Sniffles2 (ref. 82) (v.2.6.0), which was run in each regular and mosaic modes to seize each dominant and subclonal variants. The outputs from these two modes have been mixed and any duplicate variants have been eliminated, in addition to any variants mapping to contigs aside from the E. faecium chromosome and variants smaller than 100 nucleotides. The ensuing per pattern structural variant counts are proven in Prolonged Knowledge Fig. 9f (Supplementary Desk 6).
To find out whether or not ISL3 variants continued or turned fastened throughout sequential isolates from the identical affected person, we carried out longitudinal variant calling: for every affected person, the E. faecium chromosome from the pattern collected on the earliest time level was chosen because the reference. Reads from every time level have been mapped again to the reference metagenome meeting and handed as enter to Sniffles2 as above (Supplementary Desk 6). To search for structural variants brought on by IS component transpositions, all recognized structural variants mapping to the E. faecium chromosome have been screened to incorporate solely insertions or deletions within the 750–3,000 bp dimension vary, as this corresponds roughly to the size of IS parts. For every candidate variant, the anticipated sequence was extracted and analysed with ISEScan to verify the presence of IS parts. The ensuing IS variant calls have been compiled, and patterns of persistence or fixation have been examined by evaluating the frequencies and presence or absence of those IS-containing variants throughout every affected person’s time factors.
p11568 ISL3 structural variant neighbourhood protection
To measure the altering inhabitants frequency of the 2 recognized ISL3 structural variants within the 4 sequenced p11568 samples over time, ONT reads from every of p11568 t1–t4 have been aligned towards the t4 E. faecium contig with minimap2 (ref. 83) with the flags -x asm20, -c and –eqx. Per base learn depth was then counted for every place with samtools84 (v.1.21) utilizing the command: samtools depth -a -Q 10 -b t4.mattress {pattern}_sorted.bam > {pattern}.protection.txt. Poor high quality alignments are eliminated by the -Q 10 flag. First, the imply alignment depth throughout all the t4 E. faecium contig for every pattern was calculated. Then, relative depth per base was calculated by dividing the depth at every place by the imply depth per pattern. The t4 E. faecium contig was annotated with Bakta as described above and the neighbourhood of every ISL3 structural variant was visualized.
E. faecium isolation from stool
A single punch biopsy of stool samples from the primary (t1) and final (t4) time factors from p11568 have been suspended in 300 µl PBS. Every pattern was serially diluted tenfold into 900 µl PBS to 1:10,000. Then, 30 µl of the 1:10,000 dilutions have been plated onto BHI agar and incubated in a single day at 37 °C. The subsequent day, six colonies from every plate have been chosen and recognized by biotyping with a MALDI-TOF primarily based Bruker Biotyper per producer directions. Every recognized E. faecium colony was streaked to isolation on a BHI agar plate. To substantiate the insertion genotypes, overnights of isolates have been ready in 3 ml BHI broth and incubated with shaking at 37 °C. Overnights have been diluted 1:1,000 into 10-ml subcultures and grown to an optical density of roughly 1.0. Cell pellets have been then ready based on Plasmidsaurus ‘Normal Bacterial Genome with DNA extraction’ specs. The ensuing genomes have been annotated with Bakta as described above and the dearth of each ISL3 insertions in t1, and the presence of each ISL3 insertions in t4 have been confirmed within the assemblies.
RNA-seq of E. faecium isolates
From glycerol shares of each p11568 t1 and t4 confirmed to have the insertional genotypes (Fig. 5), 4 in a single day cultures have been ready in BHI as earlier than. Every in a single day tradition was diluted 1:100 into 5 ml of BHI and subcultured for round 3 h till OD600 reached 0.4; 5-ml microtubes (AXYGEN, catalogue no. MCT-500-C) have been ready with 0.5 ml of quenching resolution. The quenching resolution contained 90% (v/v) 100% ethanol and 10% (v/v) saturated phenol for RNA extraction (pH 4–5). Microtubes have been saved at −20 °C earlier than quenching occurred. After OD600 reached 0.4, 4 ml of every tradition was promptly quenched. The microtube was vortexed till the answer was nicely blended. All microtubes have been centrifuged at 10,000g for five min at 4 °C. The supernatant was poured off. The pellet was then lysed in 260 µl of a ready mastermix containing 25:1 PBS:10 mg ml−1 lysozyme (Sigma-Aldrich, catalogue no. L3790-10X1ML). Every pellet was resuspended in lysis buffer after which incubated at 37 °C for 30 min whereas rocking on a BenchRocker. Then, 30 µl of 20% SDS was added adopted by one other spherical of incubation at 37 °C. Then, 1.5 ml of Trizol was added and blended nicely by pipette. All subsequent steps have been carried out at room temperature. The answer was incubated for 10 min. Then, 0.5 ml of chloroform was added and ten vigorous full inversions of the tube have been made earlier than 3 min of incubation. After incubation, every tube was centrifuged at 12,000g for 10 min. Fastidiously avoiding any non-aqueous layer, roughly 800 µl of the aqueous section was transferred to a clear 2-ml Eppendorf tube. An equal quantity of 100% ethanol was added and blended by pipette till emulsions have been not seen. RNA was then purified utilizing the Zymo RNA Clear & Concentrator-25 equipment (catalogue no. R1018) based on producer directions. RNA was eluted in 25 µl nuclease-free water. RNA focus was measured by NanoDrop and high quality was confirmed by Bioanalyzer with the RNA 6000 Pico Package (Agilent; catalogue no. 5067-1513). RNA was then saved at −80 °C and sequenced with Novogene’s Prokaryotic Whole RNA sequencing service (NovaSeq PE150, 3G uncooked information per pattern).
Analyses of RNA-seq information
Uncooked RNA-seq reads have been trimmed with Trim Galore, obtainable at https://github.com/FelixKrueger/TrimGalore/, utilizing the command trim_galore –quality 30 –length 60 –paired {pattern}1.fq.gz {pattern}_2.fq.gz –retain_unpaired. Bowtie2 (ref. 85) (v.2.5.4) was used to align processed reads towards the p11568 t4 isolate reference genome with default parameters. On the premise of a poor alignment fraction (20.99%) in contrast with the opposite replicates (minimal: 97.63%), the fourth replicate from the t1 isolate (t1_4) was excluded. To get the per base depth of reads, samtools depth was used as earlier than besides no high quality threshold was applied. This meant reads mapping equally nicely to a number of areas have been distributed randomly. Moreover, a function depend matrix filtering alignments towards the t4 isolate reference genome with mapping high quality < 10 was generated with bedtools (v.2.27.1) multicov utilizing the command bedtools multicov -p -q 10 -bams {sample1}_sorted.bam {sample2}_sorted.bam. The ensuing function depend matrix, together with the Bakta annotation of the t4 isolate genome, have been used to carry out differential expression evaluation with the DESeq2 (ref. 86) bundle (v.1.42.1) in R. Solely options with at the least one depend in each t1 (n = 3) and t4 (n = 4) have been thought-about for differential expression. The adjusted P worth and log2(fold change) for every of the two,668 ensuing options have been visualized as a volcano plot. Thresholds of 1 × 10−25 and |log2(fold change)| ≥ 2 have been used to outline differentially expressed genes.
E. faecium progress and antibiotic dose–response assays
E. faecium isolates have been grown as described above. Antibiotic susceptibility testing was carried out as described beforehand46, with some modifications. In short, every pressure was grown in a single day (roughly 18 h) to stationary section in BHI medium at 37 °C. In a single day cultures have been then diluted to OD600 = 0.001 in Mueller-Hinton broth (Sigma-Aldrich) and incubated with twofold serial dilutions of TMP (Sigma-Aldrich) and SMX (Fisher) on the indicated concentrations. A ratio of 1:20 (TMP:SMX) was maintained for all remedies. Folic acid (Fisher) and folinic acid (Sigma-Aldrich) have been supplemented at 1 µg ml−1 the place indicated. Cultures have been incubated at 37 °C with steady shaking in technical duplicate (150 μl quantity per nicely) in a flat backside 96-well microtiter plate (Corning) and sealed with a Breathe-Straightforward membrane (Diversified Biotech). Absorbance (600 nm) was monitored each 10 min for twenty-four h utilizing an Agilent BioTek Synergy H1 microplate reader. Progress was calculated because the p.c space underneath the curve (AUC) relative to a management with no antibiotics added utilizing the R bundle gcplyr87.
Knowledge visualization
Preliminary plots have been generated by customized Python (Plotly) and R (ggplot) code. Plots have been organized and refined with AffinityDesigner v.2.5.7.
Scripting
Customized bash, Python, R and Nextflow scripts referenced within the strategies part can be found within the GitHub repository (‘Code Availability’). All scripts have been run in R v.4.3.2. Python scripts have been run utilizing v.3.10 with polars v.1.20.
Reporting abstract
Additional info on analysis design is offered within the Nature Portfolio Reporting Abstract linked to this text.
