Open data
Every entry in our identification keys, with its record behind it
749 field-identifiable entries in one file. Each one is what a person would actually try to name outdoors — a spider on a wall, a mineral in a hand sample, a fossil in a road cut — and each is matched, where a match honestly exists, to the public database that keeps the record for it. Free to download, free to reuse, CC BY 4.0.
What this actually is
Orvik publishes 23 free identification keys. Each one is a hand-built list of the things a person is realistically trying to tell apart in one situation — the spiders you find in a house, the berries you find on a hedge, the minerals you find in a gravel bar — with the diagnostic character that separates each from its look-alikes. That list is editorial work and it is the part that took the time.
This dataset is the other half: for every one of those entries, the identifier that lets you look it up somewhere authoritative, and the numbers attached to it there. Of the 749 entries, 570 resolved against at least one database. 502 carry a GBIF taxon key. The 420 of those that resolve to a single species stand on 634,096,345 georeferenced occurrence records between them, spanning 212 families and 74 orders, the oldest collected in 1877; the other 82 resolve to a genus, family or order, and their counts are reported per row but never added into that total, because a genus count already contains its species. 482 carry an iNaturalist taxon id. 48 are minerals with a published formula and Mohs range; 20 are fossil groups with a first and last appearance in millions of years.
The thing that does not exist anywhere else is the join. Occurrence databases are organised by taxonomy, which is the correct way to organise them and the wrong way to use them in a field: nobody standing over an unfamiliar spider knows to look under Theridiidae. A field key is organised by what you can see, and it usually has no identifiers in it at all. Putting the two in one table means a question like "of the things I could plausibly be looking at right now, which are actually recorded around here, and in this month" has an answer that can be computed instead of guessed.
What is in each row
One row per entry per key, so an animal that appears in two keys — a species in both the track key and the scat key, say — appears twice, once with each key's own name for it. That is deliberate: the row describes an entry in a key, not a species in the abstract.
- tool, entry, scientific_name — which key it belongs to, the name we give it, and the scientific name as written for a reader.
- matched_name, rank, family, order — what the database says the accepted name is. Where these disagree with the previous column, the previous column is the vernacular usage and this one is the current taxonomy.
- gbif_taxon_key, inat_taxon_id — stable identifiers. These are the columns to join on.
- gbif_records — occurrences carrying coordinates. Records without coordinates are excluded, which is why this number is lower than the total GBIF holds.
- month_01 … month_12 — the same records split by month of collection. This is the phenology column and the reason the file is worth having.
- gbif_first_year, gbif_last_year, gbif_top_country — the span of the record and where most of it comes from.
- mineral_formula, mohs_min, mohs_max, crystal_system, lustre — populated for mineral and gemstone entries only.
- fossil_first_ma … fossil_extant — populated for fossil entries only. A blank last-appearance with extant = yes means the group still has living members.
How it was built, and what it cannot tell you
The scientific name written on each key entry is parsed down to the first binomial and submitted to the GBIF name matcher. An abbreviated genus — the L. hesperus form that reads perfectly to a person — is rejected rather than guessed at, because it resolves to the wrong taxon often enough to matter. Where the match succeeds, the taxon key is used to pull a facetted occurrence count. iNaturalist is queried separately and a result is discarded if the returned genus does not match the one asked for. Minerals go to Macrostrat by name, fossils to the Paleobiology Database by taxon.
The single most important caveat is that an occurrence count measures recording effort at least as much as it measures the organism. Across this whole corpus, 15% of all records fall in May and 5% in December — a spread that says far more about when people go outdoors with a camera than about anything living. Any conclusion of the form "this species is commoner than that one" drawn straight from the raw column is probably a conclusion about observers. On the key pages we divide each species' month curve by the baseline for its own key, which removes the shared seasonal effect and leaves the part that belongs to the species; the raw columns are published so anyone can do that differently.
Two further limits worth stating plainly. Geographic coverage is heavily skewed toward North America and Europe, so a low count for a tropical species usually means nobody uploaded it rather than that it is rare. And a key entry that names a group rather than a species — "widow spiders", "oak" — is matched at whatever rank the name resolves to, so its record count aggregates everything below it and is not comparable with a species-level row. The rank column is there so you can filter those out.
Coverage, key by key
18 of the 23 keys have a database behind them. The remaining 5 — animals, birds, rocks, clouds, point types — have none, and the rows are published anyway with the identifier columns empty. There is no public occurrence database for a cloud formation, a knapped stone point or a hand-sample rock type, and inventing a number for those would be worse than leaving the cell blank.
| Key | Group | Entries | Matched | Records behind them |
|---|---|---|---|---|
| Whose egg is this? | Birds | 34 | 34 | 702,798,118 |
| What animal made this track? | Animals | 36 | 36 | 78,024,909 |
| What caterpillar is this? | Insects & spiders | 34 | 34 | 32,259,780 |
| What wildflower is this? | Plants & trees | 30 | 30 | 27,304,040 |
| What beetle is this? | Insects & spiders | 28 | 26 | 20,557,627 |
| Is this plant poisonous to dogs or cats? | Plants & trees | 73 | 69 | 14,472,693 |
| What berry is this, and can you eat it? | Food & foraging | 30 | 28 | 14,059,996 |
| Is that a butterfly or a moth, and which one? | Insects & spiders | 28 | 28 | 11,709,865 |
| Identify a tree by its leaf | Plants & trees | 43 | 43 | 5,847,119 |
| Identify a tree by its bark | Plants & trees | 34 | 34 | 5,266,559 |
| What shell did I find on the beach? | Sky, coast & ground | 29 | 28 | 3,912,270 |
| What mushroom is this? | Food & foraging | 28 | 26 | 3,699,241 |
| What succulent is this? | Plants & trees | 28 | 27 | 1,694,948 |
| What spider is this? | Insects & spiders | 31 | 31 | 1,187,533 |
| What snake did I just see? | Animals | 28 | 28 | 1,074,617 |
| Whose droppings are these? | Animals | 30 | 0 | — |
| Which bird did this feather come from? | Birds | 26 | 0 | — |
| What kind of rock is this? | Rocks & minerals | 33 | 0 | — |
| Which mineral is this? | Rocks & minerals | 36 | 28 | — |
| What gemstone is this? | Rocks & minerals | 27 | 20 | — |
| What fossil is this? | Rocks & minerals | 26 | 20 | — |
| What cloud is that? | Sky, coast & ground | 29 | 0 | — |
| What kind of arrowhead did I find? | Collecting | 28 | 0 | — |
Licence, attribution and how to cite it
The compilation — the key structure, the entry list, the join, the derived columns — is published under CC BY 4.0. Use it commercially, modify it, redistribute it; credit Orvik with a link. The underlying records belong to the four sources below and carry their own terms, which travel with any redistribution of the numbers:
- GBIF — Global Biodiversity Information Facility — CC BY 4.0 (aggregate counts). Occurrence records with coordinates, queried through the GBIF occurrence API.
- iNaturalist — CC BY-NC 4.0 (observation metadata). Research-grade observations only — an identification agreed by two or more people.
- Macrostrat mineral definitions — CC BY 4.0. Formula, Mohs hardness, crystal system and lustre.
- The Paleobiology Database — CC BY 4.0. Fossil occurrence counts and first/last appearance in millions of years.
A suggested citation: Orvik Open Field Data, snapshot 2026-09-04. Vast Flow, LLP. https://orvik.app/data/ — with the source databases credited alongside, since the occurrence counts are theirs and not ours.
The snapshot is regenerated rather than edited. If a number here disagrees with the source database today, the source database is right and this file is old; the date at the top of every page says which day it was taken. Nothing in it is typed by hand, and the build fails if anything ever is.
Reproducing it
Nothing here is a private extract. Every figure comes from a public API that anyone can call without a key or an account: GBIF's occurrence search with a taxon key and a month facet, iNaturalist's histogram endpoint restricted to research-grade observations, Macrostrat's mineral definitions, the Paleobiology Database's taxon endpoint. If you want to rebuild the file from scratch rather than trust ours, the four calls are the whole method, and the columns above say exactly which field of each response ends up where.
The one piece of judgement in the pipeline worth knowing about is the name matching, and it is deliberately conservative. An abbreviated genus is refused rather than guessed. A match that lands outside the lineage asked about is discarded — without that rule a single key entry resolved to an entire plant phylum and arrived carrying 540 million records, which is the kind of error that is invisible in a table and catastrophic in an average. A group-level row is kept but flagged by rank, never silently added to a species total.
What people do with it
The obvious use is the one the key pages already make: sort a shortlist by how much evidence stands behind each candidate, so a filter that leaves two possibilities can say which of the two is actually reported in that month. Prior probability is the single most useful thing a field key normally lacks, and it is exactly what an occurrence count is.
Beyond that: joining gbif_taxon_key against a regional occurrence download turns any of these keys into a local one, which is the version a naturalist actually wants. The month columns support a "what should I be looking for this week" list without any modelling at all. And for anyone building an identification model, the entry list is a ready-made set of confusable groups — the pairs that a key has to separate are the pairs a classifier gets wrong, and they are already grouped here by the character that separates them.
If you use it, a link back is the whole price. If a match here is wrong — and with 570 automated name matches some will be — mail the address on our contact page and it gets fixed in the next snapshot.
The keys this came out of
All 23 are free, work without an account, and show every entry with scripting off.