Open research data · PATSTAT 2026 Spring · release 20260818

DCI Patent Atlas

A patent corpus for Data Center Interconnect: the coherent optical links that connect data centres, mapped onto the DCI value chain and its citation network, and joined to a country panel that asks what this capability converts into.

Data Center Interconnect (DCI) carries traffic between data centres over coherent optical links. This project assembles the patent record behind that technology, places every application on the DCI value chain, reconstructs how the field cites itself, and then asks the economic question: does holding a position on this chain translate into national advantage in AI science and AI technology? Everything needed to reproduce the result — data, code, queries, report and paper — is published here.

1,033,230Applications
566,053Core tier
1980–2025Priority years
2,188,730Citation edges
157Applicant countries
86Patent offices

What the corpus contains

The corpus is assembled from five retrieval routes and then graded, not filtered. Classification codes, being the 53 IPC and CPC items of the reference framework; 150 keywords matched with word boundaries; a curated applicant dictionary of 943 entities; the citation neighbours of everything retrieved; and targeted recollection of the queries that hit the export limit. The routes are merged, deduplicated at three levels — application, DOCDB family and organisation — and every record keeps its provenance, so any route can be switched off and the analysis rerun.

Grading replaces deletion. Evidence from the classification, keyword and firm routes is combined into four tiers, and the tier is a column rather than a filter, which makes the sensitivity of every result to the DCI definition checkable.

Table 1. Corpus by tier. Unique applications, no duplicate identifiers.
TierApplicationsShareMeaning
Core566,05354.8%DCI in the narrow, physical-layer sense
Related94,9879.2%adjacent evidence, one route only
Peripheral360,19134.9%broad or generic matches
Excluded11,9991.2%flagged false positives, retained for audit
Master1,033,230100%analysis-ready single table

Grading shifts the country composition in a way that validates the definition. The United States falls from 48.2% of the full master to 39.8% of the core tier, while Japan rises from 14.3 to 17.3%, China from 7.3 to 10.3% and Korea from 7.0 to 9.0%. Software-adjacent broad filings drop to the peripheral tier and optical component and module filings remain, which is what a physical-layer definition of DCI should do.

Two corrections, documented rather than buried. PATSTAT Online truncates exports at 10,000 rows without warning, so every query is validated by comparing the saved archive's row count against the reported count, and split by priority year when it fails. And the first citation tally double-counted each pair by reading it from both the citing and the cited side; the figure published here is on a unique-pair basis and supersedes the earlier one. A third correction concerns keywords: substring matching produced a 27% false-positive rate in the validation sample, cut to 0.1% by word-boundary matching.
DCI patent applications by earliest priority year, 1980 to 2025
Figure 1. Applications by earliest priority year. The dashed grey line marks the annual mean. Filings rise through the late 1990s, plateau across the 2000s and rise again from the mid-2010s; the fall at the end is publication lag, not a decline in filing.

Where the patents are

The United States holds 316,255 applications by applicant country, ahead of Japan, China and Korea. The gap between filings at the Chinese office (283,568) and applications with a Chinese applicant country (48,227) reflects both domestic filing behaviour and the weaker country coverage of Chinese person records. Country attribution is not missing at random, so country-level results are checked against the complete-country subsample.

Table 2. Leading applicant countries and patent offices, full master.
Applicant countryApplicationsPatent officeApplications
United States316,255US434,862
Japan91,762CN283,568
China48,227JP109,924
Korea45,031WO70,387
Germany28,259EP48,303
Total (157 countries)—Total (86 offices)1,033,230
Applications by leading applicant country across five-year windows
Figure 2. Leading applicant countries across five-year windows, against the all-country mean.

Position on the value chain

Each application is placed on the DCI value chain by two independent routes. Classification codes decide position from the technology recorded in IPC and CPC; the applicant dictionary decides position from the filer's known place in the industry. Membership is multiple and the two routes overlap, so totals are unions rather than sums. Adjacent technologies, previously folded into the six production positions, are now carried as a seventh position of their own.

Why design and architecture needs the applicant route. System and interconnect design is not expressed in the classification scheme, so the reference table assigns it no codes at all. Its 104,309 applications, filed by Broadcom, Marvell, NVIDIA, Intel and Huawei among others, are visible only through applicants; classification alone would show this layer as empty. For the same reason the firm route is never used when ranking firms within a position: the dictionary assigns the value chain to the firm, so using it in both directions would be circular.
Table 3. Value-chain positions, seven positions, full master. Multiple memberships permitted.
PositionClassificationApplicantBothUnion
S0 Materials and substrate61,11529,2868,11182,290
S1 Design and architecture0104,3090104,309
S2 Optical core components184,833185,46261,714308,581
S3 Module assembly104,561110,27122,206192,626
S4 Advanced packaging (CPO, SiPh)119,89254,34411,793162,443
S5 Systems, submarine, operations265,149258,493112,292411,350
ADJ Adjacent technologies44,011233,9118,466182,042
Unique applications in scope758,654
Value-chain positions: totals by identification route and trajectories over time
Figure 3. Value-chain positions. a, totals by identification route. b, trajectories across five-year windows, against the mean across positions.

Firm concentration within a position is pronounced and consistent with the industry: two or three firms hold a large share of each position's patents. Consolidating acquisition lineages matters for reading these ranks. Corning's seven filing entities amount to 6,077 applications where the largest single entity alone shows 3,163, and the published firm table folds Coherent (II-VI, Finisar), Lumentum (Oclaro, NeoPhotonics, JDSU), Nokia (Infinera, Coriant, Alcatel-Lucent), Marvell (Inphi, Celestial AI), Cisco (Acacia), Broadcom (Avago) and NVIDIA (Mellanox) in the same way.

The upstream is almost invisible

Linking the trade codes of the reference framework to the corpus through their classification codes exposes a sharp asymmetry. The two transceiver headings stand on large patent stocks, while the materials chokepoint that dominates supply-chain discussion barely registers: gallium, germanium and indium link to 136 applications, doped wafers to 1,557. Control over these inputs is exercised through capacity and export licensing, not through patents, which is precisely why patent data alone cannot see it.

Table 4. Applications linked to HS headings through the classification concordance. These are patent counts, not trade volumes. Counts are on the classification seed layer and have not yet been recomputed on the graded master; the ranking, not the level, is the informative part.
HS headingTraded goodLinked applications
8517.62Transmission and switching equipment118,075
8517.79Parts of 8517 apparatus97,453
8541.49Laser diodes, photodiodes, APDs35,983
9001.10Optical fibres and bundles23,276
8542.31Processors and controllers15,135
3818.00Doped wafers (InP, GaAs)1,557
8112.92Gallium, germanium, indium136

How the field cites itself

The citation layer resolves 2,188,730 directed edges over 927,481 applications across 2,358 country pairs, on a unique-pair basis. Read as knowledge flow, the asymmetries are sharp. The United States is the largest node on both sides, but 709,708 of its citations are American applications citing American applications; setting self-citations aside changes the picture. On cross-border edges Japan is the largest net provider of cited knowledge at +98,437, followed by Canada at +29,762 and Korea at +11,976, while the United States is the largest net absorber at −90,096 and China the second at −38,873. The two deficits are of different kinds. The American one arises because the corpus is dense in American filings that draw on Japanese component patents; the Chinese one arises from a thin stock of cited Chinese filings against heavy citing of American ones, 31,485 citations in that direction alone, which is the signature of a follower position in the optical layer.

Projected onto the value chain, the direction of flow confirms the chain. Optical components and module assembly feed upward into systems, and systems is where knowledge accumulates, with 371,872 internal citations, the largest single cell of the matrix. Advanced packaging is small in volume but connects in both directions to materials, components, modules and systems, which is the junction the industry expects co-packaged optics to become.

Citation flows between the ten most active applicant countries
Figure 4. Citation flows between the ten most active applicant countries, on a logarithmic scale, with outgoing citations per country.

What the capability buys: the national panel

The corpus is the input to a second question. Does a position on the DCI chain convert into national advantage in AI? The panel joins three sources at country by field by year: DCI patents from this corpus, AI science from 2,701,689 Web of Science papers across 253 fields with fractional counting, and AI technology from 2,330,553 AI patents across 15 canonical fields. The 44,651 applications that appear in both the DCI and the AI patent masters are removed from the outcome, so no application sits on both sides of an equation.

Revealed comparative advantage is Balassa's index; relatedness follows the proximity-and-density construction of the economic-complexity literature; the DCI side enters as a cross density, the projection of a country's value-chain holdings onto the field space. Three margins are estimated separately with country, field and year fixed effects and country-clustered errors, at lags of one, three and five years: entry, meaning gaining advantage; sustain, meaning keeping it; and exit, meaning losing it.

Simulated effect of raising DCI cross density from 0 to 100 percent on entry, sustain and exit
Figure 5. Simulated probabilities of entry, sustain and exit as DCI cross density is raised from 0 to 100%, holding the other covariates at their sample means, with 95% confidence bands. Sustain rises and exit falls across the whole range; entry moves only where a scientific base is already present.
Estimated effect of each value-chain position on the three margins
Figure 6. Effect of each value-chain position, estimated one position at a time, on the three margins. Advanced packaging is the only position significant on both technology entry and science sustain.

The effects appear with a lag of one to five years, which turns the vulnerability logic of the interdependence literature into an investment timetable: capability bought after a bottleneck has already bitten arrives late by exactly that lag. Method, full tables and figures are in the panel paper, and the specification search behind them, 1,005 specifications, is published so that the exploratory-then-confirmatory sequence is visible.

Reproducing this

The published dataset is the output of collection and grading; every figure and table above is regenerated by the notebooks from those files.

git clone https://github.com/deep1003/dci-patent-atlas.git
cd dci-patent-atlas
pip install -r requirements.txt
jupyter lab notebooks/dci_patent_atlas_analysis.ipynb        # corpus, value chain, citations
jupyter lab notebooks/dci_science_relatedness_panel.ipynb    # RCA panel, regressions, figures

The notebooks are organised as tasks with steps underneath, one operation per cell. Collection from PATSTAT is documented rather than executed, since it needs credentials; the SQL and the collector are published in sql/ and scripts/. A note on method: retrieving all fields in one statement is refused by the PATSTAT cost estimator under the 2026 Spring edition, and a reduced variant returns two records in 352 seconds. Splitting the retrieval into four set-based statements joined locally returns 850 applications in 0.87 seconds.

Contents

Data note. The published file omits titles, abstracts and person names to stay within repository size limits; it retains identifiers, dates, offices, families, countries, full IPC and CPC codes, value-chain flags and the DCI tier. Keyword-based counts, which need the text fields, ship as pre-computed tables. Full records can be rebuilt from PATSTAT with the published SQL.