Files and Data Types
This page explains what kinds of inputs ipyrad2 works with, what files it writes, and how those files connect across the workflow.
What “Data Types” Means in ipyrad2
In ipyrad2, you do not need to describe details of your genomic library construction. Most necessary information is learned from the data itself. The more practical questions that a user needs to decide only affect where they will start in the pipeline:
- Are the reads pooled or already split by sample? [start with
demuxvstrim] - Do they need trimming? [run
trimor proceed tomap] - Is there a reference genome, or do you need a denovo pseudoreference? [run
denovoor not] - Is the project RAD-only, or does it include WGS BAMs that should be analyzed inside RAD-defined loci? [use
--rad-bams,--wgs-bams, or both duringassembly]
The File Model
The easiest way to understand ipyrad2 files is to think in four layers.
1. Input FASTQs
The starting point is to use either demux on pooled FASTQ data to demultiplex reads to individual samples by their unique barcodes or indices, or, if your data are already demultiplexed, to proceed directly to trim to trim, filter, and prepare your FASTQ data for read mapping.
These steps include the automatic detection of paired read files, when present, and of restriction cut site motifs on reads, which are used to identify barcodes/indices in demux, and trimmed from reads during trim.
The goal: filtered and trimmed FASTQ reads or read pairs files for each sample that can be input to denovo and/or map.
2. Reference FASTA
A reference genome is used as the anchor for locus assemblies in ipyrad2. You can either select the highest quality reference genome within your study clade, if one is available, or create a pseudo-reference genome from your FASTQ reads using ipyrad2 denovo. The latter is often useful even when a reference is available when you are assembling loci for highly divergent taxa, as it uses a graph-based splitting algorithm to identify and retain duplicated genes as distinct loci.
The goal: a FASTA reference genome that can be input to assemble.
3. Aligned reads (BAMs)
FASTQ reads or read pairs are aligned to the reference genome in ipyrad2 map using bwa-mem2 and minimally filtered by samtools to write sorted BAM files. This step generates a rich stats summary of the read alignment that is intended to help with selecting filtering parameters in the assembly step. Alignments are minimally filtered to remove QC fails, improper pairs, pairs tha do not map to the same scaffold, and .... Duplicates can be optionally removed by UMI tags (for RAD-type data) or coordinate position (for WGS data). Alignments are coordinate-sorted to be ready for downstream variant-calling.
The goal: coordinate-sorted BAM alignments that can be input to assemble.
4. Assembly Products
The main assembly step of ipyrad2 (assemble) requires as input a number of aligned read files (BAMs) and a reference genome file (FASTA).
The main products generated by assemble include:
- HDF5 database containing assembled loci, variants, and their reference coordinates.
- VCF file with filtered variant calls.
- BED file with the delimited positions of loci in reference coordinates.
- A human-readable loci.gz file for viewing alignments at each locus.
The HDF5 database is the primary input to the tools wex, lex, and snpex which are used to filter samples and sequences and write outputs in a variety of formats for downstream analyses.
The HDF5 database can also serve as direct input to a number of analysis tools that are built in to ipyrad2, either as direct impementations, or wrappers around external tools. This includes tools like PCA, sNMF, DAPC, ADMIXTURE, and popgen tools.
Example File Types
FASTQ
FASTQ is the standard text format for sequencing reads. Each read occupies four lines: a header beginning with @, the nucleotide sequence, a + separator, and one encoded quality character for every nucleotide in the sequence. Paired-end data are stored in separate Read 1 and Read 2 files whose read headers identify matching pairs.
Below is an example Read 1 FASTQ file. Your data will look different, including the read headers.
These reads have already been demultiplexed from a pooled lane by their i7 tag (AGTCGCTT), and from a pooled library by their inline barcode (already removed by demux). They retain the restriction cutsite motif (ATCGG) at the 5' end of each read, which will be detected and trimmed by trim.
This example also includes optional i5 tags that were ligated as unique molecular identifiers (UMIs) during 3RAD library preparation (for example, GAGCATGG, AGCAGGGA, GTGTCCAT, and TCTAGATT). These can be selected during trim and map to remove duplicates, when present.
@NB551405:60:H7T2GAFXY:1:11101:21625:19713 1:N:0:AGTCGCTT+GAGCATGG
ATCGGAAGGTACCAAAAGGCTCATGCCCTTGATCTATTATGAAATATGGAGTCTAGTAAAACTAGTAATGAATATATTTTAGAAAGTAAAACTTGATTCAAGCTATACAGTAGCCTGAAAAAATTAAAACTTCAAACATTTAC
+
EEEEEEEEEEEEEEEEEEEEEEEEEEEAEEEEEEEEEEEEEAEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEAEEEEE<AEEEEEEEEEEEEEEEAAEEEEEEEEEEEEAEEAEEEEEEEEEEAEEEAA<E<<<AAAAEA
@NB551405:60:H7T2GAFXY:1:11101:5344:19725 1:N:0:AGTCGCTT+AGCAGGGA
ATCGGCTGCCGTCCATCAACACCAGTCATCGGCCCACCGCATCTGCTTCCAGCCGCCAGCCGCCAGGTATTCTCTGTTCTCTTCTCATTCCATTTTTTGTGGTATAAGTTTAGGGTTCTATGTTTTAAGTATTTCGAAACCTG
+
EEEEEEEEEEEEEEEEEE6EAEEEEEEEEEEEEEEAEEEE/EEEEEEEEEAEEEEEE/EEEEEEEEEEAEEEAEEEEEAEEEEEEEAEEEEEEEEEEEEEEEEEEEEEEEEEEEEEE<AAAEEEEEEEEEEEEEEE<E/EEA<
@NB551405:60:H7T2GAFXY:1:11101:4160:19749 1:N:0:AGTCGCTT+GTGTCCAT
ATCGGGCTATCATCAACCACTACTTCTCGAACATTCACAATTGAAGGATGATTCAAGGACACAAGAGTATTGATTTCCCTCAAGTAATAAATAGGAAACCCCTCTCTTTGGTCACCCAACTTCATCTTCTTTAAGGCAACAAT
+
EEEEEEEEEEEEEEE6EEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEAEEEEEEEEEEEEEEEEEEEEEEEA/EEEEEEEEEEEEEEAAEAAE/EEEEEEEEAAEAEEEAE<A<EEA/
@NB551405:60:H7T2GAFXY:1:11101:21279:19779 1:N:0:AGTCGCTT+TCTAGATT
ATCGGCGGAGGTGATTTTCCTTCCTCTTCGGAAAGAGCTAACCATGCCTGCTTCTCAAGATCAGGATGATCATCGATTATCACTCCTTCCCATTCAAATATAGCACCTAACCATCCACAACCCATCCGTTCTAACCTAAGTA
+
EEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEE6EEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEAEEEEEEEEEAAAA/<EEEAAEEEEEAEE/AEAEEEAE
Reference FASTA
FASTA stores one or more named nucleotide sequences. Each record begins with a header line starting with >, followed by one or more lines of sequence. For a reference genome, records usually represent chromosomes or scaffolds. The first whitespace-delimited value in each header is the sequence name used in BAM, BED, VCF, and HDF5 coordinates, so it must be unique.
The example below is shortened for display. Line wrapping does not change the sequence: software joins all sequence lines until the next > header.
>A_tuberculatus_Chr01
GTTTTGATTACTTGGTGAACATAATTCACCCATAATCATAACACATGAAATGTGTAAATTTTTTATGAGGTATCAAAAATTCAAAAAGGCTGGCCCTCCT
ATCATTGCTTCTTTTACATATAGGTTTGTAGGATTTCTTTACAGTTTGAATGTAACTCAAAATTTAACTTAATGGTGTATATTAATTTAAAAAGAACTCC
CACTAACTCGGTGATGGGGTCATTTCTATAGCTAAAAGTATAATCAAATCCTTAAAAAAATTTCACTATTGAGTATGTGATGATTATTGTGGTGGTCGCT
TAAATTCTGTGTGCAATGAGAGGATAGTATATTAATTGTTGGTGGAATTTGTCCCTGCAGAAAATATTTTATAATTGATACATCTTATCATCTCATATTC
CTTGGACTTTGAACATGCAATTGATTTTGGCCAGGTTGTTCGAGAAATACAATACTCGAGGAGCGGCGTTACCGTAAAAACAGAGGATGGTTCCATCTAC
GAGGCTAATTACTTAATTTTGTCTGTTAGCATTGGTGTACTTCAAAGTGATCTAATTTCTTTTCGTCCACCTTTACCCGTATGTATTCTTCTTTATATTT
...
>A_tuberculatus_Chr02
...
The reference passed to map must be the same reference passed to assemble. Indexing tools create companion files for the FASTA, including a .fai index that records sequence names and lengths.
Pseudoreference FASTA
ipyrad2 denovo writes denovo_reference.fa when a suitable external reference is unavailable or when a data-derived reference is preferable. It is a valid FASTA file, but its records are inferred RAD loci rather than chromosomes. Headers such as locus_1_1 and locus_1_2 identify separately retained sequences from the same graph component. This allows the denovo workflow to preserve supported duplicated or paralogous copies as distinct mapping targets.
When paired-read arm structure is preserved, ipyrad2 joins the two arms with exactly 50 N characters. The spacer represents unknown intervening sequence, not observed sequence. Single-end loci and successfully merged paired reads do not require it.
This shortened example is taken from an ipyrad2 denovo reference. The line containing only N characters is the 50-base spacer.
>locus_1_1
TCATTCACCCTTCTGGTGCATTATCTTCTCATTTTACTCGATAACTTGGGTCT
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
ACTCAAATGAGTACTCCTTGTTCAAAGATATCCATTTGTATGTCTATCATATA
>locus_1_2
TCATTCACCCTTCTGGCGCATTATCTTCTCATTTGACTCGATAACTCGGGTCT
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
CAAATGAGTACTCCTTGTTCAAAGATATCCATTTGTACTTCTATCATATATAT
The pseudoreference is used by map like an external reference FASTA. It is an intermediate mapping reference, not the final assembled loci or HDF5 database.
BAM
ipyrad2 map writes one coordinate-sorted BAM alignment per sample. BAM is a compressed binary representation of SAM, so inspect it with samtools rather than opening it as text. Each BAM is accompanied by an index, normally sample.filtered.bam.bai, which lets programs retrieve selected intervals without scanning the whole file.
For example, samtools view -h sample.filtered.bam renders BAM records as SAM text. The long sequence, quality, and optional tag fields below are shortened for readability.
@HD VN:1.6 SO:coordinate
@SQ SN:locus_1_1 LN:562
@SQ SN:locus_2_1 LN:336
LH00150:...:UMI_GCCTTGTT 99 locus_2_1 1 8 21M2D28M9I16M2D63M = 198 336 TTCTCACC... IIIIIIII... NM:i:18 AS:i:72 RG:Z:bella-JJ85-plate_J2
The header states that the file is coordinate sorted (@HD) and describes reference sequences (@SQ). An alignment record contains, in order:
- read name (
QNAME) - bitwise flags (
FLAG) - reference sequence and 1-based position (
RNAME,POS) - mapping quality and alignment operations (
MAPQ,CIGAR) - mate reference, mate position, and template length (
RNEXT,PNEXT,TLEN) - read sequence and base qualities (
SEQ,QUAL) - optional typed tags such as edit distance (
NM), alignment score (AS), and read group (RG)
Here, flag 99 describes the first read of a properly paired alignment, = means its mate maps to the same reference, and the CIGAR string records matched bases, deletions, and insertions. assemble uses these alignments and coordinates to determine read support within each locus.
VCF (.vcf.gz)
assemble writes filtered SNP calls to NAME.vcf.gz, a block-gzip-compressed Variant Call Format file, together with a CSI index (NAME.vcf.gz.csi). Inspect it with bcftools view NAME.vcf.gz or another VCF-aware program. The final VCF is restricted to retained loci, contains SNP rather than indel records, and keeps sites that pass the final site filter.
The example below, shortened to three samples, comes from a recent assembly. Fields are tab-delimited in the file.
##fileformat=VCFv4.2
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT alaschanica-DE237 axillaris-DE37 axillaris-JJ125
locus_9_12 13 . C T 98.3163 PASS DP=585;AC=0;AN=10;MQ=39;F_MISSING=0.642857;AF=0;MAF=0 GT:DP:AD ./.:0:0,0 ./.:4:4,0 ./.:1:1,0
The fixed columns give the reference sequence (CHROM), 1-based position (POS), reference and alternate alleles (REF, ALT), variant quality (QUAL), filter status, and site annotations (INFO). Here the INFO fields include total depth (DP), alternate allele count (AC), called allele count (AN), mapping quality (MQ), missing-data frequency, allele frequency (AF), and minor allele frequency (MAF).
FORMAT defines the sample columns. In GT:DP:AD, GT is the diploid genotype, DP is sample depth, and AD gives reference and alternate read depths. Typical genotypes are 0/0 (reference homozygote), 0/1 (heterozygote), 1/1 (alternate homozygote), and ./. (missing). A VCF can retain an alternate allele in ALT even when genotype masking leaves no called copies among retained samples, as in this example (AC=0).
LOCI (.loci.gz)
assemble writes NAME.loci.gz, a gzip-compressed, human-readable collection of final multiple-sequence alignments. Each locus contains the reference row (assembly_reference_sequence) followed by empirical samples retained at that locus. All rows within a locus have the same aligned length.
This compact example is schematic but preserves the exact record structure:
assembly_reference_sequence ACGTACGTCCGTAACGTA
sample_A ACGTACGTCCGTAACGTA
sample_B ACGTACGTTCGTAACGTA
sample_C ACGTACGTTCGTATCGNN
// * - |0:locus_16_1:1-18
A, C, G, and T are unambiguous consensus bases. IUPAC codes such as R, Y, or S represent heterozygous calls, N represents missing or masked sequence, and - within an alignment row is a gap. Samples removed by final sample-row filters are omitted from that locus in this text file.
The // line terminates the locus. Its marker string aligns with sequence columns: * marks a parsimony-informative variable site, - marks another variable site, and spaces mark constant sites. The suffix is index:scaffold:start-end, where the locus index is zero-based and the displayed reference interval is 1-based and inclusive.
BED (.bed)
assemble writes NAME.bed to map every retained locus to the assembly reference. It is a four-column, tab-delimited BED file with no header.
locus_9_12 0 328 6
locus_13_2 38 242 6
locus_16_1 0 328 13
locus_17_1 0 328 13
The columns are:
- reference scaffold or pseudoreference locus name
- 0-based start coordinate, included in the interval
- 0-based end coordinate, excluded from the interval
- number of empirical sample sequences retained at that locus
The first row therefore spans 328 bases (328 - 0) on locus_9_12 and contains sequence data from six samples. The synthetic assembly_reference_sequence row is not counted in column four. These 0-based, half-open intervals are also used by the coordinate maps inside the HDF5 database.
HDF5 (.hdf5)
assemble writes NAME.hdf5, the structured binary database used by ipyrad2 export and analysis tools. HDF5 is not a line-oriented text format. It stores named, typed, multidimensional datasets and metadata attributes, allowing tools to retrieve selected samples, loci, genomic windows, or SNPs without reading the whole assembly.
You can inspect it with h5py. This code reports the datasets in a recent 14-sample assembly:
import h5py
with h5py.File("assembly.hdf5", "r") as database:
print(sorted(database.attrs))
for name in sorted(database):
print(name, database[name].shape, database[name].dtype)
['names', 'nsnps', 'reference', 'scaffold_lengths', 'scaffold_names', 'version']
genos (14, 56208, 3) uint8
phy (15, 516518) uint8
phymap (3071, 5) uint32
reference (56208,) uint8
snpsmap (56208, 5) uint32
The core datasets are:
| Dataset | Axes or columns | Information stored |
|---|---|---|
phy |
samples (including reference) x aligned sites | Concatenated locus alignments stored as ASCII-encoded nucleotide and IUPAC characters. |
phymap |
loci x (scaff, phy0, phy1, pos0, pos1) |
Maps each phy slice to a 0-based, half-open reference interval. scaff indexes the root scaffold_names attribute. |
genos |
empirical samples x SNPs x 3 | The two VCF allele indexes followed by the ASCII-encoded nucleotide or IUPAC genotype character. Missing allele indexes are 255 and their character is N. |
reference |
SNPs | ASCII-encoded reference allele at each SNP. This differs from the root reference attribute, which stores the reference FASTA path. |
snpsmap |
SNPs x (loc, loc_idx, loc_pos, scaff, pos) |
Maps each SNP to its locus, within-locus position, reference scaffold, and genomic position. All values are 0-based. |
sample_dp |
empirical samples x SNPs | Per-sample VCF depth (FORMAT/DP) in current assemblies. |
site_qual |
SNPs | VCF site quality (QUAL) in current assemblies. |
The example database predates sample_dp and site_qual, so they are absent from its listing; current ipyrad2 assemblies write both. Root attributes store sample names, schema version, SNP count, reference path, and scaffold names and lengths. Dataset attributes store column names, indexing conventions, and sample order where needed.
An HDF5 from assemble contains both sequence-backed and SNP-backed data. An SNP HDF5 created from an external VCF with ipyrad2 vcf2hdf5 lacks phy and phymap, so sequence-based tools such as wex and lex require the assembly form.
Exporting from HDF5
The main assembled database written by assemble is a sequence-backed HDF5 file. It is the hub for many downstream operations because it can support sequence-aware workflows and also serve SNP-based analyses after SNP datasets are written into it.
Some downstream tools only need SNP data. Those SNP-capable HDF5 inputs can come from two places:
- directly from
assemble - from
ipyrad2 analysis vcf-to-hdf5when the starting point is an external VCF
From there, users can either stay inside ipyrad2 and run analyses directly, or export data in other forms:
wex: window-based sequence exportslex: locus-based sequence exportssnpex: SNP-centered exports
That means the workflow is no longer only about reaching a final text file. Users can stop at structured assembled data and branch into several downstream paths from the same dataset.
Mixed RAD/WGS Workflows
This is one of the most powerful additions in ipyrad2.
The current assemble workflow accepts:
- RAD BAMs through
-d/--rad-bams - optional WGS BAMs through
-w/--wgs-bams
The important rule is that RAD BAMs still define the shared loci. WGS BAMs do not create the locus set on their own. Instead, they are normalized, filtered, and analyzed within those RAD-defined loci.
That is useful because it lets one project combine two kinds of information:
- RAD samples define homologous reduced-representation windows across the dataset
- WGS samples contribute coverage and genotype information inside those same windows
This is a real expansion of what ipyrad can do, and it can be very powerful in mixed sampling designs. But it is still not the same thing as general whole-genome assembly or unrestricted shotgun locus discovery. The WGS support is specifically WGS-in-RAD-loci support.
Assembly-Focused Command Flow
| Command | Primary input | Primary output | Notes |
|---|---|---|---|
demux |
pooled FASTQ plus barcode/index metadata | sample FASTQ files | Skip this step if reads already arrive split by sample. |
trim |
sample FASTQ files | trimmed FASTQ files plus trim reports | Uses a fastp-based trimming path with cutsite-motif awareness. |
denovo |
trimmed reads | pseudoreference FASTA and related outputs | Use when no suitable external reference is available. |
map |
trimmed FASTQ plus reference or pseudoreference FASTA | coordinate-sorted BAM files | These BAMs become the main input to assemble. |
assemble |
RAD BAMs, optional WGS BAMs, reference FASTA | NAME.hdf5, NAME.vcf.gz, NAME.loci, NAME.bed, stats |
RAD BAMs define loci; optional WGS BAMs are analyzed within those loci. |
Downstream Export and Analysis Flow
| Command | Primary input | Primary output | Notes |
|---|---|---|---|
analysis wex / lex / snpex |
assembled HDF5 | window, locus, or SNP exports | Export tools, not assembly steps. |
analysis vcf-to-hdf5 |
external VCF | SNP-capable HDF5 | Lets downstream SNP analyses start from external variant calls. |
analysis pca, snmf, dapc, admixture |
SNP-capable HDF5 | method-specific result tables | These consume filtered SNP datasets directly. |
analysis popgen |
assembled HDF5 or SNP-capable HDF5 | population-genetic summary tables | Available statistics depend on whether sequence-backed data, SNP-backed data, or both are present. |
Common Workflow Shapes
- RAD-only workflow: pooled FASTQ ->
demux->trim->map->assemble-> export or analysis - Mixed RAD/WGS workflow: RAD FASTQ ->
trim->map-> RAD BAMs, plus external WGS BAMs ->assemblewith RAD BAMs defining loci and WGS BAMs evaluated inside those loci - External VCF workflow: VCF ->
analysis vcf-to-hdf5-> SNP-capable HDF5 -> clustering or PCA-family analyses
Common Misunderstandings
- Do not confuse the older ipyrad library-type taxonomy with the current ipyrad2 interface. The taxonomy is still useful for describing data, but the workflow is now driven by subcommands and file transitions.
- Do not treat the assembled HDF5 as a generic archive. In ipyrad2 it is an active structured analysis product.
- Do not assume every analysis consumes the same kind of HDF5. Some methods need only SNP data, while others depend on sequence-backed assembled data.
- Do not over-interpret the WGS addition. It is powerful because it brings WGS samples into RAD-defined loci, not because ipyrad2 has become a general-purpose whole-genome assembly framework.
Where to Go Next
- Read What Is ipyrad? for the short project overview.
- Read The ipyrad Ethos for the project design goals.
- Read Installation if you still need to set up the environment.
- Read Quick Guide for the end-to-end workflow.
- Read trim and map for the main data-preparation steps.
- Read the Analysis Guide when you are ready to work from assembled or SNP-capable HDF5 outputs.