Licensed Datasets For Foundation AI
AI Oil builds a trusted European licensing framework for proprietary knowledge
We prepare licensed datasets for advanced language and multimodal Foundation AI systems.
Working with institutions, publishers, enterprises and public organizations, we transform proprietary knowledge into legally usable datasets for Foundation AI.
Our partners unlock new long-term licensing revenue while retaining full control over their intellectual property.
Across Europe, valuable scientific, institutional and industrial knowledge remains inaccessible through public web datasets and conventional web crawling.
Every dataset is delivered with documented provenance, licensing records and compliance metadata suitable for enterprise AI development.
The Offline Data Gap
European Knowledge Assets
From Knowledge To Licensed AI Datasets
Polish Court Rulings Corpus v1.0
Public records, engineered. Our full pipeline — cleaning, structuring, deduplication, residual-PII masking — demonstrated on Poland's open judicial data.
- / 01
Section-level structure: rulings split into operative part and reasoning, ready for summarization and legal-reasoning tasks.
- / 02
QC layer that detects and masks residual personal data which survived official court anonymization.
- / 03
Deduplicated with MinHash; template near-duplicates removed, full provenance preserved.
- / 04
Public domain under Art. 4 of the Polish Copyright Act — official documents carry no copyright. Delivered with a full datasheet.
3.35 MB · ZIP: JSONL sample + datasheet + license · stratified random sample (seed=42)
We'll email you the download link (64 MB, gzipped JSONL). No spam.
Checksums (SHA-256)
aioil_korpus_pl_sample_500.zip c80e76ae27ec82f50afad094235d37c8dcf88ccda5642eedfc3f1bf1515c9dbf aioil_korpus_pl_sample_10k.zip 581ece029bb2863f75eeb0318ad0d300f7f9559ec130e8377f0faad7b001ef8d
CONTACT
AI OIL works with organizations that own valuable knowledge assets as well as companies developing Foundation AI systems.
If your organization is exploring opportunities to license proprietary knowledge or acquire licensed datasets for Foundation AI, we would be pleased to discuss potential collaboration.