Skip to content

Open data

Every entry in our identification keys, with its record behind it

749 field-identifiable entries in one file. Each one is what a person would actually try to name outdoors — a spider on a wall, a mineral in a hand sample, a fossil in a road cut — and each is matched, where a match honestly exists, to the public database that keeps the record for it. Free to download, free to reuse, CC BY 4.0.

Entries 749 Matched 570 Records behind them 634,096,345 Snapshot 2026-09-04 Licence CC BY 4.0

What this actually is

Orvik publishes 23 free identification keys. Each one is a hand-built list of the things a person is realistically trying to tell apart in one situation — the spiders you find in a house, the berries you find on a hedge, the minerals you find in a gravel bar — with the diagnostic character that separates each from its look-alikes. That list is editorial work and it is the part that took the time.

This dataset is the other half: for every one of those entries, the identifier that lets you look it up somewhere authoritative, and the numbers attached to it there. Of the 749 entries, 570 resolved against at least one database. 502 carry a GBIF taxon key. The 420 of those that resolve to a single species stand on 634,096,345 georeferenced occurrence records between them, spanning 212 families and 74 orders, the oldest collected in 1877; the other 82 resolve to a genus, family or order, and their counts are reported per row but never added into that total, because a genus count already contains its species. 482 carry an iNaturalist taxon id. 48 are minerals with a published formula and Mohs range; 20 are fossil groups with a first and last appearance in millions of years.

The thing that does not exist anywhere else is the join. Occurrence databases are organised by taxonomy, which is the correct way to organise them and the wrong way to use them in a field: nobody standing over an unfamiliar spider knows to look under Theridiidae. A field key is organised by what you can see, and it usually has no identifiers in it at all. Putting the two in one table means a question like "of the things I could plausibly be looking at right now, which are actually recorded around here, and in this month" has an answer that can be computed instead of guessed.

What is in each row

One row per entry per key, so an animal that appears in two keys — a species in both the track key and the scat key, say — appears twice, once with each key's own name for it. That is deliberate: the row describes an entry in a key, not a species in the abstract.

  • tool, entry, scientific_name — which key it belongs to, the name we give it, and the scientific name as written for a reader.
  • matched_name, rank, family, order — what the database says the accepted name is. Where these disagree with the previous column, the previous column is the vernacular usage and this one is the current taxonomy.
  • gbif_taxon_key, inat_taxon_id — stable identifiers. These are the columns to join on.
  • gbif_records — occurrences carrying coordinates. Records without coordinates are excluded, which is why this number is lower than the total GBIF holds.
  • month_01 … month_12 — the same records split by month of collection. This is the phenology column and the reason the file is worth having.
  • gbif_first_year, gbif_last_year, gbif_top_country — the span of the record and where most of it comes from.
  • mineral_formula, mohs_min, mohs_max, crystal_system, lustre — populated for mineral and gemstone entries only.
  • fossil_first_ma … fossil_extant — populated for fossil entries only. A blank last-appearance with extant = yes means the group still has living members.

How it was built, and what it cannot tell you

The scientific name written on each key entry is parsed down to the first binomial and submitted to the GBIF name matcher. An abbreviated genus — the L. hesperus form that reads perfectly to a person — is rejected rather than guessed at, because it resolves to the wrong taxon often enough to matter. Where the match succeeds, the taxon key is used to pull a facetted occurrence count. iNaturalist is queried separately and a result is discarded if the returned genus does not match the one asked for. Minerals go to Macrostrat by name, fossils to the Paleobiology Database by taxon.

The single most important caveat is that an occurrence count measures recording effort at least as much as it measures the organism. Across this whole corpus, 15% of all records fall in May and 5% in December — a spread that says far more about when people go outdoors with a camera than about anything living. Any conclusion of the form "this species is commoner than that one" drawn straight from the raw column is probably a conclusion about observers. On the key pages we divide each species' month curve by the baseline for its own key, which removes the shared seasonal effect and leaves the part that belongs to the species; the raw columns are published so anyone can do that differently.

Two further limits worth stating plainly. Geographic coverage is heavily skewed toward North America and Europe, so a low count for a tropical species usually means nobody uploaded it rather than that it is rare. And a key entry that names a group rather than a species — "widow spiders", "oak" — is matched at whatever rank the name resolves to, so its record count aggregates everything below it and is not comparable with a species-level row. The rank column is there so you can filter those out.

Coverage, key by key

18 of the 23 keys have a database behind them. The remaining 5 — animals, birds, rocks, clouds, point types — have none, and the rows are published anyway with the identifier columns empty. There is no public occurrence database for a cloud formation, a knapped stone point or a hand-sample rock type, and inventing a number for those would be worse than leaving the cell blank.

Coverage per identification key, snapshot 2026-09-04. "Matched" counts entries that resolved against at least one external database.
Key Group Entries Matched Records behind them
Whose egg is this? Birds 34 34 702,798,118
What animal made this track? Animals 36 36 78,024,909
What caterpillar is this? Insects & spiders 34 34 32,259,780
What wildflower is this? Plants & trees 30 30 27,304,040
What beetle is this? Insects & spiders 28 26 20,557,627
Is this plant poisonous to dogs or cats? Plants & trees 73 69 14,472,693
What berry is this, and can you eat it? Food & foraging 30 28 14,059,996
Is that a butterfly or a moth, and which one? Insects & spiders 28 28 11,709,865
Identify a tree by its leaf Plants & trees 43 43 5,847,119
Identify a tree by its bark Plants & trees 34 34 5,266,559
What shell did I find on the beach? Sky, coast & ground 29 28 3,912,270
What mushroom is this? Food & foraging 28 26 3,699,241
What succulent is this? Plants & trees 28 27 1,694,948
What spider is this? Insects & spiders 31 31 1,187,533
What snake did I just see? Animals 28 28 1,074,617
Whose droppings are these? Animals 30 0
Which bird did this feather come from? Birds 26 0
What kind of rock is this? Rocks & minerals 33 0
Which mineral is this? Rocks & minerals 36 28
What gemstone is this? Rocks & minerals 27 20
What fossil is this? Rocks & minerals 26 20
What cloud is that? Sky, coast & ground 29 0
What kind of arrowhead did I find? Collecting 28 0

Licence, attribution and how to cite it

The compilation — the key structure, the entry list, the join, the derived columns — is published under CC BY 4.0. Use it commercially, modify it, redistribute it; credit Orvik with a link. The underlying records belong to the four sources below and carry their own terms, which travel with any redistribution of the numbers:

A suggested citation: Orvik Open Field Data, snapshot 2026-09-04. Vast Flow, LLP. https://orvik.app/data/ — with the source databases credited alongside, since the occurrence counts are theirs and not ours.

The snapshot is regenerated rather than edited. If a number here disagrees with the source database today, the source database is right and this file is old; the date at the top of every page says which day it was taken. Nothing in it is typed by hand, and the build fails if anything ever is.

Reproducing it

Nothing here is a private extract. Every figure comes from a public API that anyone can call without a key or an account: GBIF's occurrence search with a taxon key and a month facet, iNaturalist's histogram endpoint restricted to research-grade observations, Macrostrat's mineral definitions, the Paleobiology Database's taxon endpoint. If you want to rebuild the file from scratch rather than trust ours, the four calls are the whole method, and the columns above say exactly which field of each response ends up where.

The one piece of judgement in the pipeline worth knowing about is the name matching, and it is deliberately conservative. An abbreviated genus is refused rather than guessed. A match that lands outside the lineage asked about is discarded — without that rule a single key entry resolved to an entire plant phylum and arrived carrying 540 million records, which is the kind of error that is invisible in a table and catastrophic in an average. A group-level row is kept but flagged by rank, never silently added to a species total.

What people do with it

The obvious use is the one the key pages already make: sort a shortlist by how much evidence stands behind each candidate, so a filter that leaves two possibilities can say which of the two is actually reported in that month. Prior probability is the single most useful thing a field key normally lacks, and it is exactly what an occurrence count is.

Beyond that: joining gbif_taxon_key against a regional occurrence download turns any of these keys into a local one, which is the version a naturalist actually wants. The month columns support a "what should I be looking for this week" list without any modelling at all. And for anyone building an identification model, the entry list is a ready-made set of confusable groups — the pairs that a key has to separate are the pairs a classifier gets wrong, and they are already grouped here by the character that separates them.

If you use it, a link back is the whole price. If a match here is wrong — and with 570 automated name matches some will be — mail the address on our contact page and it gets fixed in the next snapshot.

The keys this came out of

All 23 are free, work without an account, and show every entry with scripting off.

Browse the free tools