> ** Work in Progress (WIP) Notice** This platform and accompanying research are part of an ongoing, open-source initiative by the Gram Disha Trust and IIT Delhi. We are publishing these early baselines to invite collaborative critique, vernacular corrections, and field validation from farmers, agronomists, and legal scholars. The tools shall evolve continuously with community input.*
“If the technology decreases the operational (input) cost and complexity of the smallholder – consider it – else send it back to the drawing board”
- Gram Disha Trust AgriTech Maxim

Does a seed or plant variety proven to thrive few villages away rarely reaches the farmer who needs it most — not because it wouldn't grow there, but because no one ever told her it would?
Picture a smallholder farmer in a drought-prone district, preparing her land for an uncertain monsoon. She plants the same seed variety her family has sown for generations, not because it thrives, but because it is the only seed the local shop sells. Thousands of miles away, inside a government biodiversity register, sits the record of a seed bred and proven for exactly her soil, her rainfall, and her changing climate. It has an official registration number and proven drought-resistant traits. Yet, she will likely never know it to sow it.
India’s public research institutions and the Protection of Plant Varieties and Farmers’ Rights Authority (PPVFRA) do immense, invaluable work documenting these resilient crop varieties every month. There is also the National Bureau of Plant Genetic Resources and its compendium of data on native seeds and planting materials alongside numerous CSOs that have painstakingly built Peoples Biodiversity registers.
However, a structural digital divide exists between these institutional archives and grassroots smallholders. The challenge is not just data availability, but data fragmentation and schema incompatibility. Each institution and organization collects and stores knowledge in its own isolated format, ranging from static PDF gazettes and clinical botanical tables to unstructured field notes. Because there is no common interoperable schema, these rich datasets cannot communicate with one another. More critically, these technical architectures rely on esoteric laboratory codes and statutory jargon, failing to map onto the common vernacular languages and traditional names that farmers actually recognize. In many cases, an institutionally registered variety is already conserved and cultivated by local communities under a regional vernacular name, a vital biological and cultural connection that remains completely obscured by data silos.
The challenge today is no longer generating knowledge or building human connectivity, robust networks of extension officers, CSOs, and farming collectives already exist across the country.
The true opportunity lies in last-mile data liberation. Once this agronomic knowledge is liberated, made easily available then seeds and planting materials suited to hyper-local agroclimatic conditions become instantly identifiable.
Armed with that clarity, smallholders and community networks can actively create the pathways to source, exchange, and multiply those climate-resilient varieties.
This is just one use-case among many.
So, how do we translate and pipe this wealth of public knowledge directly into the hands of the farmers who need it most?
By combining open, participatory access with the latest advances in artificial intelligence, we have kick-started an innovation to solve these complex, “wicked” agricultural challenges.
At the heart of this effort remains a singular maxim – will it lower input costs and reduce operational complexity for the smallholder?
So, let’s see what VARTA-Seeds brings to the drawing board.
So What is VARTA-Seeds?
VARTA – Vital Agroecological Resource Tracking Analytics – is an open Digital Public Infrastructure (DPI) Initiative anchored by Gram Disha Trust in collaboration with academic institutions and grassroots agroecological networks. This is producing individual Digital Public Goods (DPGs) for open public usage for Agroecological Production systems.
In recent months, two foundational pillars of this ecosystem – VARTA-IoT (for microclimate sensing) and VARTA-AI (for hyperlocalized agroecological reasoning) – were published for general public availability.
Continuing with these efforts the next DPG being shared here is VARTA-Seeds.
Its evolving definition, presently is –
VARTA-Seeds is an open-domain Digital Public Good that puts institutional seed biodiversity data directly into the hands of grassroots farmers. Using localised AI and vector search, it translates dense regulatory gazettes into accessible vernacular formats mapped to specific agroclimatic zones. Operating as a participatory commons, the platform accepts public, field-level data inputs, allowing its vector memory to continuously learn, grow, and evolve from crowdsourced farmer logs to strengthen regional seed sovereignty.
This innovative module is being developed in close collaboration with Core Stack team at IIT Delhi, by M.Tech Computer Science scholars – Shreshth Verma and Khushi Deshmukh – to develop this module. The technical architecture and agroecological framework are being jointly mentored by Prof. Aaditeshwar Seth (IIT Delhi) and Ashish Gupta-jee (Gram Disha Trust).
A core technical objective of this initiative also is establishing seamless interoperability between distinct public infrastructures. Specifically, VARTA-Seeds is being architected to integrate directly with the CoRE Stack platform, allowing seed data to enrich broader environmental and landscape-level analytics.
It is with the firm conviction that digitalisation does not need to be a top-down, capital-intensive monolith thrust upon rural communities.
It is entirely feasible to build sovereign infrastructure from the grassroots upwards linking one focused Digital Public Good (DPG) at a time until a robust, resilient DPI emerges organically.
While comprehensive technical specifications shall be detailed in subsequent publications, the indicative architecture visualised below demonstrates this approach. It illustrates how multiple independent, self-governed DPI systems can dynamically interact with one another through open protocols – proving that complex digital ecosystems can be built democratically from the ground up.

Simple Information on Seeds should not require a law degree to interpret
We started this project on a simple premise data that could help a farmer survive a bad monsoon should not be trapped behind complex language and rigid document layouts.
While a seed variety’s suitability to heat, drought, or salinity is technically a matter of public record, “public” only has meaning if the information is genuinely accessible.
While institutions such the PPVFRA regularly publishes journals detailing climate-resilient traits, Distinctness, Uniformity, and Stability (DUS) profiles, and regional adaptations, this vital knowledge remains trapped in static, complex multi-column PDF documents in Hindi and English.
So while the information is available it sits behind voluminous journals with complex tables.
It is also the case that such large datasets do not help realise the potential impact of such seeds and planting materials across India.
A smallholder in Himachal Pradesh or Maharashtra cannot reasonably download, parse, and cross-reference these massive digital archives on a basic smartphone or feature phone.
While commercial seed vendors and venture-capital-backed agritech platforms possess the capital to scrape and index this data for proprietary gain, grassroots farmers are left dependent on external intermediaries. VARTA-Seeds eliminates this dependency.
We set out to do something unglamorous but overdue – parsing thousands of dense botanical records into a direct, vernacular-first cognitive tool.
The platform overlays this institutional data onto an interactive, searchable interface organized by three vital indicators: crop type, regional adaptation, and stress tolerance.
By mapping these varieties directly onto the geography of India, users can instantly see where specific resilient seeds are registered and by whom.
Best of all, this resource is available for free, with no login, no paywall, and no gatekeeping.
As new regulatory journals are published by the authorities, our backend architecture continuously ingests and indexes the latest datasets, ensuring this digital commons grows autonomously alongside the needs of Indian farmers.
The backend system continously works on newer data sets when available and continues to build the archive.
See for yourself -
In this iteration – the parsing of the data across multiple PPVFRA journals is clarified, however it is still in the English language.
VARTA-Seeds – Agridoc Renderer – Click to open in New Tab
Climate doesn't read district boundaries - Geospatial Visualisation - WIP*

Making the archive readable was only half the challenge. A variety’s official record lists the districts where it was formally registered, usually wherever the breeder happened to run field trials. That is an accident of paperwork, not an ecological boundary of where the plant can actually grow. Two neighbouring districts can share nearly identical rainfall, heat, and soil, yet, on paper, a variety proven in one is treated as nonexistent in the other.
Here we can see the mapping overlay of multiple varieties registered in PPVFRA database journals between 2024 – 2026.
Nature does not consult a government gazette before deciding where a plant will thrive. So, we asked a different, more honest question –
Instead of “where was this variety registered?”, what if we asked “where does this variety’s biological climate profile actually match?” – Using real atmospheric and soil data rather than a paperwork trail.
Given a district's real climate, which registered varieties would genuinely suit it, including those that the official records never formally connected to that place?
To answer that, we built a second analytical layer beneath the archive – a high-resolution climate profile for every district in India.
Rather than relying on broad seasonal averages, this layer computes satellite and reanalysis weather data broken down by each crop’s actual growth stages. This precision is critical – the week a rice plant is flowering matters far more to its survival than an average temperature taken across a four-month monsoon. This data may also shift based on the agroclimatic zone. E.g. the transplant of paddy in Himachal Pradesh happens later than, say, Haryana. Therefore the matching should by Hyperlocal to region and not a general national average.
Matching a variety to a location stage by stage, rather than season by season, reflects how a farmer actually experiences agricultural risk-
Not “was this a dry year?”, but “was it dry the specific week my crop needed water most?”
It is precisely this level of hyper-local response that fits the technology suitably with the other DPGs such as VARTA-IoT.
This computational architecture driving this stage-by-stage analysis across 573 Indian districts is powered server-side via Google Earth Engine (GEE), structured through the four-step data pipeline illustrated below. This output eventually integrates with the existing module of the Core-Stack DPI.

The Important Bits - Use Cases
This is an interesting topic. While the work in still in progress and under development – this immediate parsing and visualization brings forth important use cases for such as DPI right away.
A DPG is only as valuable as its ability to solve practical, everyday problems for those on the ground. By bridging institutional botanical registries with localized climate realities and stage-by-stage weather data, VARTA-Seeds enables following distinct, actionable use cases across the agroecological ecosystem. The six scenarios below are not an exhaustive functional specification, but an open invitation. They represent starting points to spark a wider conversation among farmers, agronomists, civil society, and systems architects on how unmediated, grassroots-up Digital Public Infrastructure can ultimately serve the smallholder.
When a monsoon is delayed, heatwaves strike, or soil salinity rises, farmers need immediate access to alternative crop varieties without relying on commercial seed vendors or proprietary agritech intermediaries.
- The Practical Reality – A cultivator in the Pangna Valley or Marathwada can query the platform in their regional Indic dialect to find drought-tolerant paddies or frost-resistant pulses whose biological needs match their district’s (or rather Geo located farms’) actual agroclimatic and soil profile.
- The Impact – Farmers independently evaluate and source verified traditional landraces suited to their exact microclimate – for free, with no paywalls, gatekeeping, or commercial bias. Based on their present circumstances VARTA-IoT on a farmers field or Block level Trial Plots may even recommend real-time conditions for most suitable variety mapping.
Transforming static text into an interactive, geospatial map turns VARTA-Seeds into an evidentiary instrument for civil society, agronomists, and farming consortiums to safeguard common property resources and claim their statutory entitlements.
- The Practical Reality – Under Section 21 of India’s PPV&FR Act, the public has only a three-month window to oppose a commercial registration. By making spatial anomalies immediately visible- such as a corporation claiming a traditional acid lime line across an entire state- communities can spot questionable claims before the window closes. Furthermore, the platform demystifies Section 39, which explicitly guarantees a farmer’s legal right to save, exchange, and sell unbranded seed of protected varieties.
- The Impact – Grassroots networks can correlate regional data to trigger powerful statutory remedies: asserting Community Rights (Section 41), filing for Revocation (Section 34) against exaggerated claims, or establishing evidence of historical conservation to demand equitable Access to Benefit Sharing (ABS) from the National Gene Fund when commercial breeders utilize farmer-developed genetic resources.
Official regulatory journals regularly trap traditional seed heritage under clinical alphanumeric codes (e.g., IC-XXXXX), stripping away ancestral nomenclature and local identity.
The Practical Reality – Through the open contribution layer, seed-savers and farmers can tag official records with regional Indic names (restoring identities like Thooyamali or Kullu Rajmash) alongside field-verified morphological traits. This immutable Germination Event Ledger creates a decentralized repository of timestamped public prior art, documenting exactly where and by whom indigenous landraces have been historically cultivated and conserved. It also informs the Regulator of such updates which may be – from time to time – included as part of the official gazette.
The Impact – The decentralized Germination Event Ledger allows communities to log real-time empirical performance. When a farmer asks the system for planting advice, the AI grounds its recommendations not just in clinical DUS tables, but in immutable, real-world field logs from fellow cultivators – What companion cropping maximized yield? What was the actual germination rate during last year’s unexpected frost? This ledger serves a dual purpose. Legally, it builds an unshakeable evidentiary shield against biopiracy, preventing commercial entities from claiming traditional varieties as “novel” breeding discoveries. Agriculturally, it powers unmediated, peer-to-peer seed networks, allowing farmers across districts to directly discover, exchange, and trade verified heirloom seeds and germination insights without relying only on formal corporate seed supply chains.
True to our grassroots-up architectural philosophy, VARTA-Seeds is engineered not as a standalone silo, but as an autonomous module designed to communicate seamlessly with broader digital public infrastructures.
The Practical Reality – By integrating directly with IIT Delhi’s CoRE Stack platform, plant variety datasets can be cross-referenced with macro-level environmental telemetry, landscape analytics, and real-time microclimate data from VARTA-IoT.
The Impact – Environmental planners, public researchers, and grassroots institutions can design regional climate adaptation strategies that link seed suitability directly to soil health, watershed restoration, and long-term agroecological transitions.
For decades, digital agricultural platforms have suffered from an English-language bias, locking smallholders out of critical regulatory and botanical data. A true Digital Public Good must speak the language of the soil.
- The Practical Reality – By integrating localisation such as – Sarvam AI’s foundational Indic LLMs directly into the cognitive pipeline, VARTA-Seeds may just translate dense statutory jargon, botanical DUS (Distinctness, Uniformity, and Stability) tables, and climate parameters into natural, conversational regional dialects, whether in Hindi, Marathi, Tamil, or local Pahari phrasing.
- The Impact – Farmers and local self-help groups can interact with the archive via voice or text in their mother tongue without losing scientific accuracy. This bridges the digital divide, ensuring that linguistic sovereignty goes hand-in-hand with seed sovereignty—free from the gatekeeping and licensing costs of commercial foreign AI models.
Formal commercial seed registries demand strict Uniformity and Stability (the “U” and “S” in DUS characterization), treating a seed as an industrial widget that must perform identically across vast regions. However, community-managed seed systems thrive on genetic and phenotypic heterogeneity. Traditional landraces are dynamic, living populations that intentionally express genetic variance to survive localized environmental shocks.
- The Practical Reality – A traditional landrace like Molakolukulu (the high-starch paddy renowned for making Andhra’s GI-tagged Poothrekullu sweets) possesses built-in defense mechanisms against fungal epidemics like Rice Blast or coastal salinity. While a uniform monoculture risks total field collapse during an outbreak, individual plants within a diverse community landrace naturally vary in disease resistance and flowering timelines. Through VARTA-Seeds, smallholders can proactively map where Molakolukulu is adapted, check real-world field germination rates logged by peers in the Germination Ledger, and locate available seed stock at local Community Seed Banks (such as the Atreyapuram CSB or regional FPCs).
- The Impact – This creates a powerful dual advantage for the smallholder. Agronomically, farmers gain immediate, unmediated access to climate-resilient native seeds without relying on expensive commercial vendors. Legally, VARTA-Seeds automatically transmutes these vernacular field observations into formal MCPD V.2.1 and UPOV TGP schemas, establishing unshakeable public prior art under Sections 39 and 41 of the PPV&FR Act to permanently shield community biodiversity against corporate biopiracy.
These use-case applications only scratch the surface of what becomes possible when agricultural data is liberated from institutional siloes. Whether it is a local self-help group building localized value chains around a heritage grain, or an independent agronomist designing regional climate resilience models, the true potential of a Digital Public Infrastructure lies in what the community builds upon it.
The goal is not to prescribe every end-use, but to provide the unmediated, vernacular-first foundation required to keep that conversation, and smallholder sovereignty, moving forward. As the intent is to make this open source DPG access, with time other interesting use-cases may also be co-developed to support impactful applications.
The Open Architecture - The Technical Stack

To ensure VARTA-Seeds remains a resilient Digital Public Good, free from recurring cloud compute costs, proprietary paywalls, or black-box dependencies, the platform is engineered as a 100% self-hosted, deterministic, and open-source pipeline.
Decades of unstructured Plant Variety Journal documents are transformed into searchable, interactive digital services through a seven-stage automated process –
1. Document Ingestion – Digital copies of historical and current regulatory gazettes are gathered and placed into an automated processing queue.
2. Text Extraction – Specialized software strips away complex magazine layouts, images, and decorative formatting to extract clean, readable plain text from the documents.
3. Record Separation – Automated rules scan the extracted text to identify document boundaries and cleanly separate the data into individual, standalone crop variety records.
4. Artificial Intelligence Parsing & JSON Generation – On-premise language models – executed locally using Ollama and powered by Qwen– read each separated record to identify, understand, and organize key botanical traits, agricultural characteristics, and legal details, transmuting raw text directly into standardized, schema-validated JSON objects.
5. MongoDB Storage & API Standardization- The organized records are securely saved into a central database MongoDB. This allows for standardized API access and interoperability for future integrations. During this step, automated checks remove duplicate entries and standardize all state and district names to match official government geographical boundaries.
6. Process Management – An automated background controller supervises the entire workflow from start to finish, ensuring smooth processing, preventing errors, and safely archiving completed documents.
7. Interactive Web Services – The clean, standardized JSON data stored in MongoDB powers two accessible web tools – a search engine for users (farmers, extention officers, researchers, government officials or CSO experts etc.) to browse individual plant varieties, and an analytical dashboard displaying comprehensive charts, maps, and trait breakdowns.
[INSERT GITHUB REPOSITORY CARD / IFRAME HERE: <GitHub Codebase Link Pipeline Shreshth VARTA-Seeds to via>] Insert the live GitHub repository widget here so developers and systems architects can directly inspect, clone, or contribute to the open-source parsing scripts and orchestration architecture.
The Road Ahead
Work in Progress - Multi-Model Benchmarking & Indic Integration
Extracting complex botanical tables and clinical plant characteristics from historical documents without errors or data hallucination is a significant technical challenge. Currently, Ollama is utilized to run open-weight models locally, specifically Qwen (Qwen 2.5 14B), which serves as the core cognitive engine for structuring raw text into validated JSON objects without relying on paid cloud APIs or compromising data sovereignty.
However, artificial intelligence architectures evolve rapidly, and no single model is infallible. Active benchmarking and stress-testing of multiple open-weight large language models against Qwen is essential going forward. Continuously comparing model performance ensures the highest possible extraction accuracy, eliminates hallucinated data, and optimizes processing speed across thousands of historical records.
Furthermore, true smallholder sovereignty requires breaking the English-language barrier at the root. A critical milestone on the immediate roadmap is the testing and integration of Sarvam AI’s foundational Indic language models directly into this pipeline. While models like Qwen handle structural JSON extraction, Sarvam AI is specifically proposed to natively Indianise the archive, translating clinical statutory jargon and botanical characteristics into regional dialects. This integration will allow farmers and local self-help groups to search, read, and interact with the database using voice or text in their mother tongue. Eventually a chat or voice interface may also be suitable to deploy since browsing a database browser may also be cumbersome for smallholders. By layering Sarvam AI’s vernacular speech-to-text (STT) and conversational LLM engines over the MongoDB API, the platform will allow farmers to bypass search forms entirely. Cultivators will be able to query the archive via conversational voice messages in their mother tongue, asking natural questions about local seed availability, disease resilience, or germination rates, and receive synthesized, spoken agronomic guidance. This conversational abstraction ensures that the complex database remains invisible, reducing user complexity to simple oral communication.
Expanding the Pipeline – Building a Federated Botanical Commons
Currently, the ingestion engine focuses on scouring historical and monthly PPVFRA Plant Variety Journal PDFs. However, a true Digital Public Good cannot rely on a single statutory archive. A major architectural objective on the immediate roadmap is expanding the ingestion pipeline to federate multiple public-domain botanical and germplasm databases:
Ex-Situ Genebank Integration (NBPGR) – Expanding automated parsing to query the National Bureau of Plant Genetic Resources (NBPGR) database. Extracting standardized Passport Information (accession numbers, collector IDs, native origin) and Characterization/Evaluation data allows cross-referencing legal variety registrations directly against physical germplasm holdings conserved in national genebanks.
In-Situ Community Seed Atlases (MSSRF & CSB Networks): Integrating open-domain repositories such as the Community Seed Bank Atlas by M.S. Swaminathan Research Foundation (MSSRF) and other grassroots ledgers. This maps traditional landraces and heirloom seeds actively conserved by indigenous communities, bringing visibility to informal farmer-to-farmer seed networks that exist outside statutory intellectual property registries.
Automated Web Federation & Ontologies – Moving beyond deterministic PDF extraction requires building robust web scrapers, API connectors, and botanical ontologies. These pipelines will dynamically query web search portals, harmonize disparate agricultural naming conventions, and unify records across state agricultural universities (SAUs) and international public-domain repositories. This is also an important aspect for a DPI to achieve for overall interoperability.
Two-Way Community Synchronisation – Interoperability must flow in both directions. Rather than merely harvesting data from public seed atlases, the architecture is being designed to allow community seed banks and farmer producer organizations (FPOs) to dynamically link their living inventories and field observations with the Germination Ledger, closing the loop between national repositories and grassroots cultivators.
Automated Schema Transmutation – Bridging Grassroots Data with Global Standards
Formal institutions like NBPGR and PPVFRA rely on complex, highly structured scientific descriptor schemas, specifically the FAO/Bioversity Multi-Crop Passport Descriptors V.2.1 (MCPD V.2.1) and the International Union for the Protection of New Varieties of Plants Test Guidelines Procedures (UPOV TGP). While essential for global scientific rigor, these complex frameworks can create prohibitive data-entry barriers for public contributors, civil society, and smallholders.
When a contributor enters raw, unstructured botanical observations from the public domain, the underlying cognitive pipeline automatically transmutes the simplified input into standardized JSON objects fully compliant with MCPD V.2.1 and UPOV TGP schemas. If specific technical descriptors are missing, the system dynamically gathers and imputes the missing parameters from open domain reference sources, presenting the completed profile back to the user for final checking and verification. This automated schema transmutation ensures seamless, bidirectional scalability between informal grassroots smallholders and formal global scientific systems.
Where this is going – Building the Bridge
The VARTA-Seeds archive is live and expanding every month. Currently, climate matching is operational as a pilot specifically for rice, with the exact same analytical approach ready to extend to every other crop in the registry.
However, turning the insight that “these varieties match your district’s climate” into an actionable tool that can be used in the field is not a challenge that technology can solve in isolation.
None of this replaces what farmers already know – generations of observation about local soil, weather, and seed behavior remain the deepest source of agricultural wisdom that exists.
What is being built here is strictly a bridge –
A mechanism for the public record, paid for by the public and meant for the public, to actually reach the people for whom it was written.
This brings us back to the Maxim.
The Open Invitation – Collaborative Ground-Truth
Ultimately, this continuous engineering effort is governed by the foundational AgriTech Maxim of the Gram Disha Trust-
“If the technology decreases the operational (input) cost and complexity of the smallholder – consider it – else send it back to the drawing board.”
True appropriate technology is not measured by the architectural complexity of its artificial intelligence models, but by its capacity to hyperlocalise analytics, systematically eliminate financial and technical barriers for grassroots cultivators.
Because smallholder farming environments are dynamic and resource-constrained, VARTA-Seeds remains an active work in progress, committed to iteratively refining its pipelines to ensure that global scientific rigor always remains accessible, lightweight, and economically viable for the everyday farmer. Even thus far, the maxim holds well for the innovation thus far too.
To transform this open-source infrastructure into an everyday agricultural reality, active collaboration is needed across three vital domains –
Agronomists – To test and verify the climate-matching logic against what actually happens on the ground, ensuring recommendations reflect field realities rather than just theoretical data.
Climate Scientists – To sharpen how district-level environmental risks, heatwaves, and rainfall patterns are analyzed through a crop’s actual biological growth stages.
Farmers and Extension Workers – To evaluate whether these digital tools are genuinely useful in daily field operations, and to provide direct feedback on how the interface and advisory outputs need to evolve.
Civil Society Stakeholders – To validated, improve and localise the solution based on the lived realities of the smallholders. Going forth it will be interesting to see how other such initiatives may collaborate on such innovations.
The data has been liberated from the archives. Now, let’s build an agricultural digital public infrastructure that speaks the language of the soil.
