We make
machines see.
Professional visual data for machine learning, built from a 20-year
single-source professional imaging archive.
RAW CAPTURES
MEDIUM-FORMAT RAWS
HUMAN-EDITED TIFF MASTERS
ARCHIVE
Approximately 400,000 original professional RAW captures, including around 350,000 Hasselblad medium-format RAWs, together with approximately 250,000 historical professionally human-edited TIFF masters.
Around 70TB of proprietary professional image data, created through a controlled, colour-managed production workflow rather than scraped or aggregated from the web.
Art Image Library has developed a Google Cloud platform that links original source captures with historical production files and can generate consistent 16-bit machine baselines from the original RAW data.
This creates the potential for structured comparison between:
providing training and evaluation data based on real professional image-processing decisions rather than synthetic approximations.
MACHINE LEARNING
The archive includes EXIF and IPTC metadata, colour-reference information where available, structured lineage and associated rights information.
The platform also supports machine-generated enrichment, structured JSON and Parquet outputs, vectors, embeddings and semantic search.
The underlying collection is predominantly art and cultural imagery, while the visual-processing relationships within the data are potentially applicable across broader image-model training and evaluation tasks.
Art Image Library controls the photographic rights in the collection, while rights in underlying works vary by image, with a substantial proportion in the public domain.
Material can be segmented according to proposed use, and appropriately screened evaluation datasets can be prepared for prospective partners.
We are interested in conversations around AI training and evaluation data, data licensing, technology partnerships and strategic opportunities.