Per-project CSV splits of the Starrydata2 dataset, refreshed daily.
| Repository | Description | Update Schedule | Period |
|---|---|---|---|
| Google Drive | Latest dataset only | Twice daily at 00:00 and 12:00 | from 2024/06/13 |
| Figshare | Past datasets | Daily until 2024/06/06, then monthly | from 2022/12/22 |
| Github | Past datasets | As needed | from 2019/7/11 until 2022/12/22 |
sample_id values — the same sample_id can refer to samples from different papers. Joining tables on sample_id alone will mix unrelated data. Workaround: join on (SID, sample_id), which is unique. A corrected dataset (starrydata_dataset_renumbered.zip) is available on Google Drive. The live database will be fixed in the week beginning 2026-09-07; datasets published after that date already include the fix.links page to read directly from this repo's daily manifest.json. New projects (e.g. OrganicThermoelectricMaterials) now appear the day they're added to the DB, instead of waiting for the monthly Figshare mirror. Added totals and per-project counts (including figures) to manifest.json.ThermoelectricMaterials, BatteryMaterials, MagneticMaterials, etc.) can now be downloaded separately as papers / samples / curves files, alongside the full unsplit dataset..csv.gz) to reduce file size. Load directly with pandas.read_csv(url, compression="gzip") or decompress before opening in Excel.figure_name field to curve dataset.all_curves.csv is now starrydata_curves.csv.project_names and created_at to the paper dataset.all_samples.csv in certain applications, such as Excel, by adding a BOM."2024-05-17 00:00:01 JST+0900" to "2024-05-17 00:00:01 GMT+0900 (JST)".["299.8597", "324.8683"] → [299.8597, 324.8683]updated_at, created_at, and composition_details to all_samples.csv.