Paired Policy Commons RIS and CSV collection

A page-preserving, identifier-checked method for enriching bibliographic records with provider-supplied country and topic metadata.

Implementation 1.1 · 20 September 2026 · validated with Python 3.14.6, pandas 3.0.5 and pyarrow 25.0.1

Purpose. Every structurally valid CSV is stored beside the RIS file collected for the corresponding Policy Commons search page. RIS supplies rich bibliographic fields and abstracts. CSV adds provider-supplied country and topics. To maximise collection throughput, page-level RIS–CSV agreement is not evaluated during download. Final enrichment is performed later by persistent record identifier, never by row or page position; unmatched records are skipped at that stage.

1. Frozen collection contract

Queryartificial intelligence, logical search
Coverage1950–2026, one publication year at a time
ModulesPOCO_POGO World Governments; POCO_PCWL World Cities; POCO_POHE Public Health and Social Care; POCO_PSPL North American State, Provincial and Territorial Governments; POCO_ALL Global Think Tanks
Page size100 results
Preferred item scopeartifact_type=document:Document. Historical batches that were originally exported without this filter retain their recorded manifest contract so that the CSV reproduces the RIS page rather than rewriting collection history.
Output unitOne RIS and one CSV per module, year, partition and page
AuthenticationAn existing user-authenticated Chrome session exposed locally through the Chrome DevTools Protocol at 127.0.0.1:9223

Frozen AIGOV5 execution used for the worked result

Query IDaigov5
Exact queryArtificial Intelligence AND AI Governance AND AI Safety AND AI System AND AI Policy
ModulesPOCO_POGO, POCO_PCWL, POCO_POHE, POCO_PSPL
FiltersPublications; document:Document; gov:Government; 100 results per page
Sort coveragedate_desc pages 1–100 plus date_asc pages 1–53
Provider count15,206 displayed search hits at collection time
Frozen observations153 RIS files and 153 CSV files, 15,300 observations in each format

The scripts do not bypass access controls and do not store credentials, cookies, CSRF tokens or signed download addresses. Reproduction requires an authorised Policy Commons account and must comply with the applicable licence.

2. Architecture and state transition

  1. Read the canonical module-year inventory and resolve its relative raw-source links.
  2. Locate every valid page_NNNN.ris and its nearest manifest.jsonl.
  3. Recover the exact module, year, query, sort order, artefact filter, partition and page number used for that RIS.
  4. Open the same search contract in the authenticated browser.
  5. Submit the currently rendered result identifiers to /api/export/many/ with export_type: csv.
  6. Download the short-lived object to .csv.part.
  7. Validate CSV decoding, required schema and the cardinality of the current CSV export only.
  8. Atomically rename the structurally valid temporary file to page_NNNN.csv. Later integration evaluates record-level matches.
DISCOVER RIS → RECOVER SEARCH CONTRACT → RENDER AUTHENTICATED PAGE
    → REQUEST CSV EXPORT → DOWNLOAD TO .part → VALIDATE CSV STRUCTURE
    → ATOMIC SAVE → RECORD-LEVEL MATCH AND QA LATER

3. Search partitioning above the 10,000-result boundary

Policy Commons exposes at most 100 pages of 100 results for a single search ordering. A module-year above 10,000 results is therefore divided before export. The partition planner is deterministic relative to the provider state observed at planning time.

  1. Request the module-year count. Zero produces a recorded zero-result year. Counts at or below 10,000 use a single partition.
  2. For larger sets, recursively add mutually exclusive title-prefix clauses using a–z and 0–9.
  3. Each sibling excludes preceding sibling clauses. This prevents overlap when Policy Commons tokenisation makes prefix behaviour non-exclusive.
  4. Create a residual partition with NOT (title:a* OR ... OR title:9*).
  5. If a residual remains above the limit, use newest and oldest sort directions only when the observed count is no greater than 20,000. Publication-date subdivision is available as a fallback.
  6. Store the complete query and expected count in partition_plan.jsonl and the page manifest.
if total <= 10_000:
    partitions = [base_query]
else:
    partitions = split_large_partition(
        parent_query="artificial intelligence",
        alphabet="abcdefghijklmnopqrstuvwxyz0123456789",
        mutually_exclusive=True,
        retain_residual=True,
    )

4. Authenticated export at code level

The search results are rendered in Chrome. The code reads only provider-generated DOM identifiers and supported content types, then submits the export request inside that authenticated origin.

const results = Array.from(document.querySelectorAll('.search-item'))
  .map(item => ({
    object_id: item.dataset.id,
    content_type: item.dataset.type
  }))
  .filter(item => ['file', 'artifact', 'topic', 'organization']
    .includes(item.content_type));

const response = await fetch('/api/export/many/', {
  method: 'POST',
  credentials: 'include',
  headers: {
    'Content-Type': 'application/json',
    'X-CSRFToken': document
      .querySelector('[name=csrfmiddlewaretoken]')?.value
  },
  body: JSON.stringify({
    export_type: 'csv',
    content_type: 'search',
    object_id: 0,
    results
  })
});

The response contains a short-lived signed object address. It is used immediately through the authenticated browser request context and is never written to a manifest.

5. Page acceptance tests

TestRequired conditionFailure action
RIS structureTY count equals ER count and is non-zeroDo not request a companion for an invalid RIS
CSV decodingStrict UTF-8 with optional BOMReject temporary file
CSV schematype,title,url,publisher,year,topics,country are present; the observed export also contains people,coi,language,issn,isbnReject and retry
Export cardinalityCSV rows equal the item count submitted for the current CSV exportReject and retry because the CSV export itself is incomplete
Page cardinalityNot evaluated during collectionEvaluate during integration QA
Artefact identityNot evaluated during collectionUse later as a secondary record-level check; never join by position
Bibliographic identityAt integration, compare CSV.coi with RIS.ID after conservative normalisationJoin matching records only; skip unmatched records from the enriched output while retaining them in raw storage and QA logs
AtomicityValidation occurs on .csv.part; only an accepted payload becomes .csvDelete incomplete temporary payload
csv_info = inspect_csv(partial_csv)
assert csv_info["records"] == rendered_export_items
csv_info["pair_status"] = "not_evaluated_at_collection"
partial_csv.replace(final_csv)  # atomic rename on the same volume

On the first audited 2026 page, 100 RIS records matched 100 CSV rows in order by both COI and artefact ID. All 100 CSV rows contained country information and 96 contained topics. This is a pipeline smoke test, not a population completeness estimate.

6. Record linkage and field provenance

The later bibliographic integration must not join on title, row number or page position. The primary join is the persistent content identifier:

primary_key   = normalise(CSV.coi) == normalise(RIS.ID)
secondary_key = artifact_id(CSV.url) == artifact_id(RIS.UR)
accept_pair   = primary_key and secondary_key

published_in_country   <- CSV.country
policy_commons_topics  <- CSV.topics
ris_keywords           <- repeated RIS.KW       # retained separately
issuer_country         <- later evidence layer  # never overwritten by CSV country

country is preserved as provider-supplied publication-country evidence. It is not automatically treated as issuer country, target country, policy jurisdiction or analytical country. topics is stored independently from RIS keywords and analyst-derived classifications.

7. Deterministic Golden Set construction

The collector prioritises fast, lossless capture. The Golden Set builder then operates at record grain. It never assumes that page 39 in one export still represents page 39 in a later export because provider holdings and rankings can change between requests.

  1. Parse all frozen RIS records and retain the existing deduplicated RIS master.
  2. Concatenate all CSV observations and attach source file and source row provenance.
  3. Normalise CSV.coi and RIS.ID by lower-casing, removing DOI resolver prefixes and trimming terminal slashes.
  4. Parse the numeric artefact identifier from /artifacts/{id}/ in both URL fields.
  5. Deduplicate CSV observations on the compound key (normalised COI, artefact ID). Preserve all contributing filenames and observation counts.
  6. Accept a RIS–CSV pair only when both key components agree. Do not use title, page number, sort position, year or publisher as a substitute join.
  7. Retain unmatched RIS, unmatched CSV and within-key field conflicts in separate audit files.
  8. Merge CSV country, topics, people, issn and isbn into the paired master with field-level provenance.
  9. Create a fixed-seed 100-record human-validation sample. The paired output remains a Golden Set candidate until original-source adjudication is complete.
join_coi         = normalise(CSV.coi) == normalise(RIS.ID)
join_artifact_id = artifact_id(CSV.url) == artifact_id(RIS.UR)
accept_pair      = join_coi and join_artifact_id

assert not golden.duplicated(["join_coi", "join_artifact_id"]).any()
assert golden["join_coi"].ne("").all()
assert golden["join_artifact_id"].ne("").all()

Reproduced AIGOV5 results

CSV observations15,300
Unique CSV compound keys11,320
Deduplicated RIS records14,696
Exact RIS–CSV pairs11,306, 76.9325% of the RIS master
AI-included, document-eligible pairs11,285
Provider publication country present11,292
Provider Topics present10,192
Unmatched records3,390 RIS records and 14 unique CSV records, retained for audit
Within-key field conflicts15 compound keys, retained in csv_field_conflicts.csv
Duplicate accepted compound keys0

Interpretation boundary. golden_set_master is identifier-exact and provenance-complete for the accepted pairs. Its human_validation_status is pending. It must not be described as a fully adjudicated ground-truth set until the fixed-seed validation records have been checked against original documents.

8. Idempotence, retries and recovery

  • Before requesting a page, validate any existing CSV structurally. A valid export is skipped because matching is deferred to the record level.
  • A structurally invalid existing CSV is moved to a provenance-preserving quarantine path, not overwritten.
  • A live search with no result cards is skipped after a 20-second bounded wait so that a changed historical partition cannot block the sweep.
  • A browser-verification challenge or other transient failure receives at most two short attempts with a five-second delay.
  • Unavailable and failed exports are written to unresolved.jsonl. The main sweep continues.
  • checkpoint.json stores queue index, total, module, year, RIS path, state and UTC timestamp.
  • manifest.jsonl records row count, columns, size, SHA-256, paired RIS path, not_evaluated_at_collection status and the later join-key definition.
  • Re-running the collector is safe because structurally valid CSV files are revalidated and skipped.

9. On-disk layout

artificial_intelligence_.../
└── POCO_POGO/
    └── 2026/
        └── document_Document_6a852c3300/
            ├── page_0001.ris
            ├── page_0001.csv
            ├── page_0002.ris
            ├── page_0002.csv
            └── manifest.jsonl

csv_companion_backfill_20260920/
├── queue_summary.json
├── checkpoint.json
├── manifest.jsonl
├── errors.jsonl
├── unresolved.jsonl
└── quarantine/

golden_set_v1/
├── golden_set_master.parquet
├── golden_set_master.csv.gz
├── golden_set_document_eligible.csv.gz
├── validation_sample_100.csv
├── unmatched_ris_records.csv.gz
├── unmatched_csv_records.csv
├── csv_field_conflicts.csv
├── input_csv_manifest.csv
└── SUMMARY.json

The canonical 1950–2026 module view contains relative symbolic links to raw source folders. The backfill resolves each physical source once, preventing the same raw folder from being processed twice when it has more than one catalogue reference.

10. Reproduction procedure

  1. Use an authorised Policy Commons account and record the exact query URL, modules, filters, page size, sort direction, page number and UTC collection time.
  2. Install the pinned runtime or record equivalent versions. The worked result used Python 3.14.6, pandas 3.0.5 and pyarrow 25.0.1. The general collectors use Playwright 1.60.0.
  3. Launch Chrome with a local debugging port and sign in interactively. Do not write cookies or signed addresses to configuration files.
  4. Freeze RIS, CSV and manifests before integration. Record SHA-256 for every input file.
  5. Run the structurally bounded CSV collection. During collection require the documented schema, UTF-8 decoding, submitted-item cardinality and atomic .part handling.
  6. Place build_golden_set_v1.py in the frozen AIGOV5 collection root beside raw_csv/ and analysis_steps5_8_v4/.
  7. Run the builder exactly once, then rerun it to confirm idempotent row counts and hashes for unchanged inputs.
  8. Require zero duplicate accepted compound keys, zero blank join components, zero temporary files and 100 unique records in the validation sample.
  9. Review unmatched and conflict files before any later release. Do not delete them merely to improve the match rate.
python -m py_compile collect_partitioned_ris.py \
  collect_csv_2026_all_modules.py backfill_csv_companions.py \
  build_golden_set_v1.py

python -u collect_csv_2026_all_modules.py
python -u backfill_csv_companions.py
python -u build_golden_set_v1.py

python - <<'PY'
import pandas as pd
d = pd.read_parquet('golden_set_v1/golden_set_master.parquet')
assert len(d) == 11306
assert not d.duplicated(['join_coi','join_artifact_id']).any()
assert d.join_coi.ne('').all() and d.join_artifact_id.ne('').all()
s = pd.read_csv('golden_set_v1/validation_sample_100.csv')
assert len(s) == 100 and not s.record_id.duplicated().any()
PY

Policy Commons holdings and metadata can change. Exact reproduction therefore means executing the frozen method and retaining dated manifests, not assuming that a later query will return identical records.

11. Public code snapshot

The repository contains code and method documentation only. Licensed RIS and CSV payloads, authentication state and signed download addresses are not published.

Citation: Chun, Y. (2026). Paired Policy Commons RIS and CSV collection: authenticated export, identity validation and recovery protocol. Implementation 1.1.