Native: tier 1
A purpose-built tool extracts, probes, converts, or writes the file in a single step. Tabular, stats packages, molecular structures, documents, and neuroscience ephys — NWB native, plus vendor formats converted to NWB.
Almost any data your research produces - tables, images, sequences, signals, and simulations - across 286 file formats in every research domain.
Your research produces data in as many forms as it asks questions: a sequencer's reads, a microscope's stacks, a climate model's cubes, a survey's tables. FAIR² takes it as it comes.
Clara reads across 286 file formats spanning every research domain and turns a single submission into four peer-reviewed outputs with lifetime hosting.
Bring the files your instruments and pipelines already generate - FASTQ and VCF, OME-TIFF and plate maps, NetCDF cubes, FITS observations, NWB recordings, mass-spec mzML - or paste a dataset DOI and Clara fetches the files and harvests the metadata published with them. It does its best with whatever you produce, and its best work when you bring the context that gives those files meaning.
286 | 3 | 35 | 6 | 8 | ~ 30 |
Clara does the reading. But the first and most important decision is yours - what actually belongs in your dataset - because a FAIR² Data Article is peer-reviewed.
The files you submit should form a dataset you have deliberately selected - one you can defend in peer review as complete, scientifically meaningful, and reusable by others, not a raw dump of everything on disk.
It doesn't need to be perfect or fully documented to start: Clara works with what you have, shows you where the gaps are, and helps you get it there.
Methodology and protocols. How the data was collected, processed, and derived — the steps a reader would need to trust and reuse it.
Equipment and instruments. Make, model, and the settings that shaped your measurements.
Variables and units. What each field means, and the units it is measured in.
The more you bring, the more Clara has to work with and the less it has to guess. This is the single biggest lever on the quality of your Data Article.
Every extension is placed in one of three tiers by how deeply Clara can read it. We are precise about the difference between parsing a file and preserving it.
A purpose-built tool extracts, probes, converts, or writes the file in a single step. Tabular, stats packages, molecular structures, documents, and neuroscience ephys — NWB native, plus vendor formats converted to NWB.
Clara writes and runs a script using one of 35 vetted scientific libraries. FITS (astropy), OME-TIFF, CZI, ND2, LIF (tifffile), NetCDF, GRIB (xarray), FASTQ, FASTA (Bio), DICOM, NIfTI, AnnData, pathology slides, MD trajectories, mass spec, and Zarr.
Recognized by the classifier and preserved alongside the parsed data, with provenance intact. Includes BAM, VCF, CRAM, annotations (GFF, GTF, BED), code, and media. The honest stance: we recognize and keep these, we do not claim to parse them.
.obj
.ply
.stl
.swc
.vtk
.7z.
bz2
.gz
.tar
.xz
.zip
.arf
.evt
.fit
.fits
.fz
.pha
.rmf
.vot
.ccp4
.chk
.cif
.cml
.cube
.ent
.fchk
.hkl
.inchi
.ins
.mae
.map
.mcif
.mmcif
.mol
.mol2
.mtz
.pdb
.pdbqt
.pqr
.res
.sd
.sdf
.smi
.smiles
.wfn
.wfx
.asc
.dem
.grb
.grib
.grib2
.h5
.hdf
.hdf5
.las
.laz
.nc
.nc4
.segy
.sgy
.zarr
.bash
.bat
.c.cfg
.conf
.cpp
.go
.h
.ini
.ipynb
.java
.jl
.js
.m
.properties
.ps1
.py
.r
.rb
.rs
.sh
.sql
.toml
.ts
.yaml
.yml
.zsh
.doc
.docx
.md
.rst
.tex
.txt
.ab1
.bai
.bam
.bcf
.bed
.bigwig
.bw
.cram
.csi
.embl
.fa
.faa
.fai
.fasta
.fastq
.fna
.fq
.gb
.gbk
.gff
.gff3
.gtf
.maf
.newick
.nex
.nexus
.nwk
.phy
.phylip
.sam
.sff
.tbi
.vcf
.wig
.dbf
.fgb
.geojson
.gml
.gpkg
.kml
.kmz
.osm
.pbf
.prj
.shp
.shx
.wkt
.aac
.aif
.aiff
.avi
.bmp
.flac
.flv
.gif
.jpeg
.jpg
.m4a
.m4v
.mkv
.mov
.mp3
.mp4
.mpeg
.mpg
.ogg
.opus
.png
.svg
.tif
.tiff
.wav
.webm
.webp
.wma
.wmv
.fcs
.mgf
.mzid
.mzml
.mzxml
.pepxml
.protxml
.traml
.czi
.dv
.ims
.lif
.lsm
.mrxs
.nd2
.ndpi
.oib
.oif
.scn
.svs
.vsi
.arc
.dcd
.edr
.gro
.itp
.lammpstrj
.mdcrd
.prm
.psf
.rtf
.top
.tpr
.trr
.xtc
.xvg
.xyz
.annot
.bdf
.bval
.bvec
.cnt
.dcm
.dicom
.edf
.eeg
.fdt
.fif
.gii
.label
.mgh
.mgz
.mha
.mhd
.mnc
.nhdr
.nii
.nrrd
.set
.stc
.tck
.trk
.vhdr
.vmrk
.abf
.ncs
.nev
.nex5
.ns1
.ns2
.ns3
.ns4
.ns5
.ns6
.nse
.ntt
.nwb
.pl2
.plx
.rhd
.rhs
.smr
.smrx
.h5ad
.h5seurat
.loom
.mtx
.arrow
.csv
.db
.dta
.feather
.json
.jsonl
.log
.mat
.ndjson
.npy
.npz
.orc
.out
.parquet
.por
.raw
.rda
.rdata
.rds
.sas7bdat
.sav
.sqlite
.sqlite3
.tsv
.xls
.xlsx
.xml
Don't see your format? That isn't a dead end - Clara handles more than this list and new formats are added often. Here's what happens with an unfamiliar or preserved format.