India is the world's largest common law system. 40M+ pending cases. 25 High Courts. 14 Tribunals. 1.7M+ advocates. And until recently, all this data was locked behind two incumbents (SCC Online and Manupatra), neither of which offers API access, vector embeddings, or citation graphs.
We built Vaquill to make Indian legal data accessible as infrastructure.
What's in the box:
20M+ Court Cases - Supreme Court, all 25 High Courts, 14 Tribunals (NCLT, DRT, ITAT, CESTAT, NGT, RERA, and more). Structured metadata per case: court, bench, date, parties, sections cited, acts referenced, case type, headnotes.
23,122 Acts & Statutes - Central, State, and Regulatory Acts. Full text, searchable, cross-referenced with case law. Amendment tracking.
Citation Graph - Case-to-case citation analysis across the full corpus. Which cases followed, distinguished, or overruled which. Good law / bad law classification. 4-level nested citation linking.
Vector Embeddings - Every case embedded with Voyage AI (1024d dense) + BM25 sparse. Hybrid semantic + keyword search. Sub-500ms API responses.
Legal Translation (Anuvad) - Translate legal documents between English and 11 Indian languages: Hindi, Tamil, Telugu, Bangla, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Urdu. 22,000+ government-verified legal terms. Not Google Translate (which turns "bail" into "land" in Hindi).
OCR Pipeline - Scanned court orders, handwritten FIRs, legacy Krutidev-encoded documents. Upload an image, get structured text back.
Access formats:
- REST API (real-time search, retrieve, filter)
- Bulk export (JSON / Parquet) for model training
- Vector search via Qdrant for RAG pipelines
- Translation API for document-level or text-level legal translation
- MCP Server (vaquill-mcp on PyPI) for direct Claude integration
Who this is for:
Legal AI companies - building legal research, contract analysis, or compliance tools that need Indian jurisdiction coverage. Harvey, Legora, and others are racing into India. The data layer is the bottleneck.
AI labs and foundation model companies - Indian court judgments are public domain (no copyright on judicial decisions under Indian law). 13M+ structured, clean, legally safe training data points. Bilingual pairs across 11 languages from our translation service.
Background verification and fintech companies - court record checks for hiring, lending, and insurance. Semantic name matching via embeddings gives better accuracy than keyword search against court records.
RegTech and compliance platforms - 23K+ acts with amendment tracking and court interpretation cross-referencing. Which cases interpreted which section of which act.
Law firms and legal departments - cross-border M&A, international arbitration, India IP disputes. Your India research just got an API.
Why now:
Manupatra (one of India's two major legal databases) just signed an exclusive data partnership with Legora in January 2026. SCC Online (the other major database) is partnered with Harvey and has copyright litigation history with Thomson Reuters.
The supply of independent, technically advanced Indian legal data just got tighter. Vaquill is one of the few remaining sources that is not locked into an exclusive deal, offers vector embeddings and citation graphs (not just raw text), and is available via API.
What we're looking for:
- Data licensing partners (AI labs, legal tech companies, BGV firms)
- API integration partners (legal platforms adding India coverage)
- Translation API partners (platforms serving multilingual users)
- Feedback from anyone building in legal tech
Try the platform: vaquill.ai
Translation: anuvad.ai
MCP: pypi.org/project/vaquill-mcp
Built by a solo engineer in India. The entire infra runs on ~$300/month.


