Download
Starrydata2 data can be obtained through two routes: (1) per-sample / per-figure web download, and (2) the official Datasets snapshots (three CSV files). Both are free for commercial and non-commercial use.
4.1 Individual download from the web
Each paper card in the paper list has csv / json links. Clicking them downloads the curve data associated with that paper. Get Data and Copy to Clipboard on the Data page are also convenient for obtaining small amounts of data.
4.2 Official Datasets snapshots (recommended)
When working with the full dataset for machine learning or statistical analysis, using the snapshots published under Datasets in the global navigation is most efficient. Each folder contains three CSV files, a README.md, and a snapshot timestamp (db_snapshot.txt).
| File | Key columns | Role |
|---|---|---|
starrydata_papers.csv |
SID, DOI, URL, issued, author, title, container_title, volume, issue, page, ISSN, publisher, project_names, created_at |
Paper bibliographic information (1 row = 1 paper) |
starrydata_samples.csv |
sample_name, sample_id, composition, composition_details, SID, DOI, sample_info (JSON), created_at, updated_at |
Sample information (1 row = 1 sample). sample_info is JSON containing domain-specific attributes. |
starrydata_curves.csv |
SID, DOI, composition, sample_id, figure_id, figure_name, prop_x, prop_y, unit_x, unit_y, x (JSON array), y (JSON array), project_names, comments, created_at, updated_at |
Curve data (1 row = 1 curve). x / y are JSON array strings of equal length. |
SID (papers ↔ samples ↔ curves) and sample_id (samples ↔ curves) as the primary keys. figure_id can be used to group multiple Curves belonging to the same figure (there is no separate Figure table).
Structure of sample_info
sample_info is a JSON string in the format {descriptor: {category: "", comment: "", extracted: ""}}. Example descriptors by domain:
- Thermoelectric (ThermoelectricMaterials):
MaterialFamily,DataType,Form,FabricationProcess,ElectricalMeasurement,ThermalMeasurement,Purity,RelativeDensity,GrainSize, and others - Magnetic (MagneticMaterials):
DataType,Form,FabricationProcess,MagneticMeasurement,saturation magnetization,coercivity,remanence magnetion,magnetic field,Measurement temperature, and others - Battery (BatteryMaterials):
Cathode active material,Anode active material,Solvent 1/2/3,Solute 1/2/3,Electrolyte Additive 1/2/3,Current density,Upper/Lower voltage limit,Cell type, and others
4.3 Key properties by domain
Thermoelectric: Temperature / Z-Seebeck coefficient / Electrical resistivity / Electrical conductivity / Thermal conductivity / Power factor / ZT / Carrier mobility / Hall coefficient
Magnetic: Temperature / Magnetic field / Magnetic field strength (H) / Magnetization / Magnetization per weight / Magnetization per volume / Magnetization (Bohr)
Battery: C rate / Cycle number / Voltage / Discharge capacity / Charge capacity
4.4 Python loading example
Basic pattern for joining the three CSV files with pandas:
import json
import pandas as pd
# Read the three CSVs
papers = pd.read_csv("starrydata_papers.csv")
samples = pd.read_csv("starrydata_samples.csv")
curves = pd.read_csv("starrydata_curves.csv")
# x / y are JSON array strings → convert to list
curves["x"] = curves["x"].map(json.loads)
curves["y"] = curves["y"].map(json.loads)
# Join sample info + paper info onto Curve (using SID and sample_id)
df = (
curves
.merge(samples[["sample_id", "sample_name", "sample_info"]], on="sample_id", how="left")
.merge(papers[["SID", "title", "issued"]], on="SID", how="left")
)
# Extract only Seebeck coefficient from the thermoelectric project
seebeck = df[
df["project_names"].str.contains("ThermoelectricMaterials", na=False) &
(df["prop_y"] == "Seebeck coefficient")
]
print(seebeck[["SID", "composition", "prop_x", "unit_x", "unit_y"]].head())
To explode into long format (1 point = 1 row):
long = seebeck.explode(["x", "y"]).rename(columns={"x": "x_value", "y": "y_value"})
long.to_csv("seebeck_long.csv", index=False)
unit_x / unit_y are in SI base unit product notation (e.g., V*K^(-1)). Use regular expressions or sympy to format them for display.
4.5 Paper archive (NIMS MDR)
The paper describing the Starrydata project (STAM:Methods 2025) can be accessed and cited via two routes. The MDR (NIMS Materials Data Repository) version guarantees a persistent URL, making it preferable for citations where link stability is important.
- MDR version (persistent URL): Starrydata: from published plots to shared materials data (MDR)
- Publisher version (Open Access): STAM:Methods (2025)
4.6 License and citation
Starrydata2 data are free to use for both commercial and non-commercial purposes. If you use the data in a paper or report, the list of papers to cite is available on the How to cite page.
- Citing the project as a whole: see /cite/
- Citing a specific version: use a Figshare snapshot (with DOI; state the retrieval date)
- For individual data: also cite the original paper (CrossRef DOI)