Skip to content
DeeptechNavigator

Methodology

How This Data Is Built

What we mean by “deep tech”

A company is deep tech when its core is novel science or engineering with a hard technical moat — not a business model, a service, or a software wrapper. The test is the technical substance, judged from what the company says it does, not the sector label alone: a company listed under “technology hardware” that only trades hardware is not deep tech.

  • Counts as deep tech— novel science, engineering or frontier R&D: semiconductors and chips, quantum, biotech and synthetic biology, energy storage and battery chemistry, robotics and drone hardware, space, photonics, advanced materials, patentable hardware, or genuinely novel AI models and algorithms.
  • Does not count — digital services, marketplaces, SaaS and pure business-model plays, consumer apps, agencies, trading, consulting, clinics / fitness / healthcare services, or a thin wrapper that merely uses an existing AI or cloud API.
  • Borderline— genuine uncertainty: a hardware or manufacturing firm whose write-up doesn’t make the novelty clear, or an “AI” company where it’s unclear whether the model is novel or an API wrapper.

How companies are classified as deep-tech

Startups are drawn from official public startup registration records (over 450,000 companies — name, industry, sector, stage, website and a self-reported one-line description). We first drop clearly non-technical industries (advertising, construction, retail, real estate, food & beverage and the like), leaving roughly 217,000 candidates to assess; the rest are kept but marked not-assessed.

Each candidate’s name, industry, sector and description is then read by a large language model and assigned one of three labels — yes (clearly deep tech), maybe (borderline, or not clear from the description) or no(a business, service or software play) — plus a short technology tag for the “yes” companies. The run is deterministic (fixed settings, one label per company), so the classification is consistent and repeatable. The deep-tech and borderline companies are the prospects carried forward and shown across the platform.

How to read the labels. Yes is high-confidence deep tech, with a short technology tag (e.g. semiconductors, biotech, advanced materials, drone hardware). Maybe is worth a human look — the description wasn’t specific enough to be sure, not a “weak yes.” Nowas screened out. A blank label means an industry we deliberately didn’t assess.

This is a screening classifier, not a definitive verdict.It reads only the company’s self-reported description and sector — not its patents, products or live website — and description quality varies, so treat borderline (and low-signal) labels as candidates for review. Being AI-based it is probabilistic: runs are consistent, but the model can misread a vague or misleading description. The technology tag is a short orientation label, not a controlled taxonomy; a deeper structured taxonomy (theme / sub-theme / maturity, aligned to India’s national deep-tech themes) has been piloted but not yet applied at scale.

Sources

Startups are drawn from official public startup registration records. Patents, designs and trademarks come from the Indian Patent Office, Design Registry and Trade Marks Registry public journals. Institution rankings use the NIRF 2025 overall ranking.

Refresh cadence

Data refreshes quarterly. Every rollup on this site is rebuilt from the underlying registries at each refresh — nothing here is computed live from a live request.

Matching, not ownership

Startup-to-IP links are found by matching normalised entity names (and, where available, corroborating state/address information) between the startup registry and patent/design/trademark filings. This is name-based matching, not a legal or authoritative record of ownership — hence every cross-dataset claim on this site says "matched", never "owns", and carries a High or Medium confidence label. High confidence means the match is corroborated by a shared state/address; medium means the names matched but nothing else confirmed it. Treat all matches as candidates, not certainties.

How companies & institutes are ranked

Raw patent counts reward paperwork: a filer can top a volume leaderboard by filing thousands of applications it never pushes through examination. So we rank by quality, not quantity, using the patent registry's own status signals — whether a filing was granted, whether examination was even requested, whether it lapsed, and how deep its claims run.

Institutes are scored 0–100: quality = 40% grant rate + 35% examination-requested share + 25% claim depth (average claims per patent, capped at 12). "Examination-requested share" is the fraction of the portfolio that actually requested examination — filings left to lapse without a request pull the score down. The rate-based sorts (quality, grant rate) apply a 20-filing floorso a two-patent, 100%-granted outlier can't top the list.

Companies (non-educational filers) are ranked by portfolio size, grant rate and lapsed share, with the same 20-filing floor on the grant-rate sort. Where institutes are ranked within a sector or technology, we use a sample-adjusted grant rate— (granted + 2) ÷ (filed + 20) — which pulls small portfolios toward the base rate so a single lucky filing can't outrank a genuine research programme. Attribution is by patent assignee record; a jointly-owned patent is credited to each co-owner. All of this is precomputed at the quarterly refresh, not live.

Company and institute rankings are computed from official patent-registry status — grant rate, whether examination was requested, lapse or abandonment, and claim depth — as an objective proxy for research quality over raw filing volume. Read them as a directional starting point for your own diligence, not investment, legal or technical advice; coverage gaps and assignee-attribution errors can shift a ranking. Verify against primary sources before relying on them, and note Deeptech Navigator takes no responsibility for decisions made on these rankings.

AI-assisted analysis & enrichment

A growing share of what you see is generated or matched by AI, not taken verbatim from an official record. This includes: the plain-language Analysis on each patent (problem, approach, novelty, claim scope, enforceability); company enrichment — brand and trade names, one-liners, funding, stage and investors — matched from public web sources; the sector narratives; and the written Insights briefings.

AI-generated content is a starting point to help you grasp the gist quickly — not legal, financial or professional advice. It can be incomplete, out of date, or matched to the wrong company (false positives and negatives). Always check it against the original source and have a qualified professional verify before you rely on it. Where a field is AI-derived we label it, and Deeptech Navigator takes no responsibility for decisions made on it.

Patent data scope

Data scope: this set concentrates on patent applications filed from 2022 onward (about 85% of records), with grants and publications from 2024 onward. It also carries a smaller tail of older records — almost entirely granted patents reaching back over a decade — so pre-2022 years show granted survivors, not full filing volume. The newest filings are still in examination, so grant counts on recent years read low by design — a coverage-window timing effect, not weak inventive output.

Trend chart baseline

Chart starts at 2022 — earlier years reflect granted-only survivorship, not full filing volume.

Technology momentum (Rising / Steady / Cooling)

Each deep-dive technology carries a momentum tag — ▲ Rising, — Steady or ▼ Cooling — read from its patent-filing history, not from headlines. We compare the two most recent fully-published filing years against the earlier baseline: more than 15% above baseline reads as Rising, more than 15% below as Cooling, and anything in between as Steady.

Why the current year is left out. A patent is typically published about 18 months after it is filed, so the current calendar year is always badly under-counted — most of its filings haven’t surfaced in the public record yet. Folding that half-empty year into the calculation would make almost every technology look like it is cooling. So we exclude the current, still-publishing year from the signal entirely — it is still shown in the filing table, marked partial — and compare only years whose record is reasonably complete. The window rolls forward on its own each year, so no year is ever hard-coded.

To avoid over-reading thin data, a technology earns a Rising or Cooling tag only when it has a genuine earlier baseline and enough complete-year filings to be meaningful; otherwise it stays Steady. In practice that means we call a technology “cooling” only when its fully-published record has genuinely declined — never as an artefact of the coverage window.

Coverage gaps

Not every startup in the source registry has been classified as deep-tech yet, and not every patent has been technology-tagged. Sector pages state their coverage where it is below 70%; the "Other" sector holds patents whose technology theme has not yet been re-bucketed into one of the 14 named sectors.

What this site does not show

Funding and investor data, where shown, is AI-matched from public web sources and clearly labelled as such (see above) — not audited financials. We do not show phone numbers, emails or street addresses for startups or individuals — inventor and filer locations are shown only at city or district level, derived from the public filing address, never the street line or postcode. We never identify the patent attorney or filing agent behind any IP asset, under any circumstance.