Paired Policy Commons RIS and CSV collection
A page-preserving, identifier-checked method for enriching bibliographic records with provider-supplied country and topic metadata.
Implementation 1.1 · 20 September 2026 · validated with Python 3.14.6, pandas 3.0.5 and pyarrow 25.0.1
Purpose. Every structurally valid CSV is stored beside the RIS file collected for the corresponding Policy Commons search page. RIS supplies rich bibliographic fields and abstracts. CSV adds provider-supplied country and topics. To maximise collection throughput, page-level RIS–CSV agreement is not evaluated during download. Final enrichment is performed later by persistent record identifier, never by row or page position; unmatched records are skipped at that stage.
1. Frozen collection contract
| Query | artificial intelligence, logical search |
|---|---|
| Coverage | 1950–2026, one publication year at a time |
| Modules | POCO_POGO World Governments; POCO_PCWL World Cities; POCO_POHE Public Health and Social Care; POCO_PSPL North American State, Provincial and Territorial Governments; POCO_ALL Global Think Tanks |
| Page size | 100 results |
| Preferred item scope | artifact_type=document:Document. Historical batches that were originally exported without this filter retain their recorded manifest contract so that the CSV reproduces the RIS page rather than rewriting collection history. |
| Output unit | One RIS and one CSV per module, year, partition and page |
| Authentication | An existing user-authenticated Chrome session exposed locally through the Chrome DevTools Protocol at 127.0.0.1:9223 |
Frozen AIGOV5 execution used for the worked result
| Query ID | aigov5 |
|---|---|
| Exact query | Artificial Intelligence AND AI Governance AND AI Safety AND AI System AND AI Policy |
| Modules | POCO_POGO, POCO_PCWL, POCO_POHE, POCO_PSPL |
| Filters | Publications; document:Document; gov:Government; 100 results per page |
| Sort coverage | date_desc pages 1–100 plus date_asc pages 1–53 |
| Provider count | 15,206 displayed search hits at collection time |
| Frozen observations | 153 RIS files and 153 CSV files, 15,300 observations in each format |
The scripts do not bypass access controls and do not store credentials, cookies, CSRF tokens or signed download addresses. Reproduction requires an authorised Policy Commons account and must comply with the applicable licence.
2. Architecture and state transition
- Read the canonical module-year inventory and resolve its relative raw-source links.
- Locate every valid
page_NNNN.risand its nearestmanifest.jsonl. - Recover the exact module, year, query, sort order, artefact filter, partition and page number used for that RIS.
- Open the same search contract in the authenticated browser.
- Submit the currently rendered result identifiers to
/api/export/many/withexport_type: csv. - Download the short-lived object to
.csv.part. - Validate CSV decoding, required schema and the cardinality of the current CSV export only.
- Atomically rename the structurally valid temporary file to
page_NNNN.csv. Later integration evaluates record-level matches.
DISCOVER RIS → RECOVER SEARCH CONTRACT → RENDER AUTHENTICATED PAGE
→ REQUEST CSV EXPORT → DOWNLOAD TO .part → VALIDATE CSV STRUCTURE
→ ATOMIC SAVE → RECORD-LEVEL MATCH AND QA LATER3. Search partitioning above the 10,000-result boundary
Policy Commons exposes at most 100 pages of 100 results for a single search ordering. A module-year above 10,000 results is therefore divided before export. The partition planner is deterministic relative to the provider state observed at planning time.
- Request the module-year count. Zero produces a recorded zero-result year. Counts at or below 10,000 use a single partition.
- For larger sets, recursively add mutually exclusive title-prefix clauses using
a–zand0–9. - Each sibling excludes preceding sibling clauses. This prevents overlap when Policy Commons tokenisation makes prefix behaviour non-exclusive.
- Create a residual partition with
NOT (title:a* OR ... OR title:9*). - If a residual remains above the limit, use newest and oldest sort directions only when the observed count is no greater than 20,000. Publication-date subdivision is available as a fallback.
- Store the complete query and expected count in
partition_plan.jsonland the page manifest.
if total <= 10_000:
partitions = [base_query]
else:
partitions = split_large_partition(
parent_query="artificial intelligence",
alphabet="abcdefghijklmnopqrstuvwxyz0123456789",
mutually_exclusive=True,
retain_residual=True,
)4. Authenticated export at code level
The search results are rendered in Chrome. The code reads only provider-generated DOM identifiers and supported content types, then submits the export request inside that authenticated origin.
const results = Array.from(document.querySelectorAll('.search-item'))
.map(item => ({
object_id: item.dataset.id,
content_type: item.dataset.type
}))
.filter(item => ['file', 'artifact', 'topic', 'organization']
.includes(item.content_type));
const response = await fetch('/api/export/many/', {
method: 'POST',
credentials: 'include',
headers: {
'Content-Type': 'application/json',
'X-CSRFToken': document
.querySelector('[name=csrfmiddlewaretoken]')?.value
},
body: JSON.stringify({
export_type: 'csv',
content_type: 'search',
object_id: 0,
results
})
});
The response contains a short-lived signed object address. It is used immediately through the authenticated browser request context and is never written to a manifest.
5. Page acceptance tests
| Test | Required condition | Failure action |
|---|---|---|
| RIS structure | TY count equals ER count and is non-zero | Do not request a companion for an invalid RIS |
| CSV decoding | Strict UTF-8 with optional BOM | Reject temporary file |
| CSV schema | type,title,url,publisher,year,topics,country are present; the observed export also contains people,coi,language,issn,isbn | Reject and retry |
| Export cardinality | CSV rows equal the item count submitted for the current CSV export | Reject and retry because the CSV export itself is incomplete |
| Page cardinality | Not evaluated during collection | Evaluate during integration QA |
| Artefact identity | Not evaluated during collection | Use later as a secondary record-level check; never join by position |
| Bibliographic identity | At integration, compare CSV.coi with RIS.ID after conservative normalisation | Join matching records only; skip unmatched records from the enriched output while retaining them in raw storage and QA logs |
| Atomicity | Validation occurs on .csv.part; only an accepted payload becomes .csv | Delete incomplete temporary payload |
csv_info = inspect_csv(partial_csv)
assert csv_info["records"] == rendered_export_items
csv_info["pair_status"] = "not_evaluated_at_collection"
partial_csv.replace(final_csv) # atomic rename on the same volume
On the first audited 2026 page, 100 RIS records matched 100 CSV rows in order by both COI and artefact ID. All 100 CSV rows contained country information and 96 contained topics. This is a pipeline smoke test, not a population completeness estimate.
6. Record linkage and field provenance
The later bibliographic integration must not join on title, row number or page position. The primary join is the persistent content identifier:
primary_key = normalise(CSV.coi) == normalise(RIS.ID)
secondary_key = artifact_id(CSV.url) == artifact_id(RIS.UR)
accept_pair = primary_key and secondary_key
published_in_country <- CSV.country
policy_commons_topics <- CSV.topics
ris_keywords <- repeated RIS.KW # retained separately
issuer_country <- later evidence layer # never overwritten by CSV country
country is preserved as provider-supplied publication-country evidence. It is not automatically treated as issuer country, target country, policy jurisdiction or analytical country. topics is stored independently from RIS keywords and analyst-derived classifications.
7. Deterministic Golden Set construction
The collector prioritises fast, lossless capture. The Golden Set builder then operates at record grain. It never assumes that page 39 in one export still represents page 39 in a later export because provider holdings and rankings can change between requests.
- Parse all frozen RIS records and retain the existing deduplicated RIS master.
- Concatenate all CSV observations and attach source file and source row provenance.
- Normalise
CSV.coiandRIS.IDby lower-casing, removing DOI resolver prefixes and trimming terminal slashes. - Parse the numeric artefact identifier from
/artifacts/{id}/in both URL fields. - Deduplicate CSV observations on the compound key
(normalised COI, artefact ID). Preserve all contributing filenames and observation counts. - Accept a RIS–CSV pair only when both key components agree. Do not use title, page number, sort position, year or publisher as a substitute join.
- Retain unmatched RIS, unmatched CSV and within-key field conflicts in separate audit files.
- Merge CSV
country,topics,people,issnandisbninto the paired master with field-level provenance. - Create a fixed-seed 100-record human-validation sample. The paired output remains a Golden Set candidate until original-source adjudication is complete.
join_coi = normalise(CSV.coi) == normalise(RIS.ID)
join_artifact_id = artifact_id(CSV.url) == artifact_id(RIS.UR)
accept_pair = join_coi and join_artifact_id
assert not golden.duplicated(["join_coi", "join_artifact_id"]).any()
assert golden["join_coi"].ne("").all()
assert golden["join_artifact_id"].ne("").all()
Reproduced AIGOV5 results
| CSV observations | 15,300 |
|---|---|
| Unique CSV compound keys | 11,320 |
| Deduplicated RIS records | 14,696 |
| Exact RIS–CSV pairs | 11,306, 76.9325% of the RIS master |
| AI-included, document-eligible pairs | 11,285 |
| Provider publication country present | 11,292 |
| Provider Topics present | 10,192 |
| Unmatched records | 3,390 RIS records and 14 unique CSV records, retained for audit |
| Within-key field conflicts | 15 compound keys, retained in csv_field_conflicts.csv |
| Duplicate accepted compound keys | 0 |
Interpretation boundary. golden_set_master is identifier-exact and provenance-complete for the accepted pairs. Its human_validation_status is pending. It must not be described as a fully adjudicated ground-truth set until the fixed-seed validation records have been checked against original documents.
8. Idempotence, retries and recovery
- Before requesting a page, validate any existing CSV structurally. A valid export is skipped because matching is deferred to the record level.
- A structurally invalid existing CSV is moved to a provenance-preserving quarantine path, not overwritten.
- A live search with no result cards is skipped after a 20-second bounded wait so that a changed historical partition cannot block the sweep.
- A browser-verification challenge or other transient failure receives at most two short attempts with a five-second delay.
- Unavailable and failed exports are written to
unresolved.jsonl. The main sweep continues. checkpoint.jsonstores queue index, total, module, year, RIS path, state and UTC timestamp.manifest.jsonlrecords row count, columns, size, SHA-256, paired RIS path,not_evaluated_at_collectionstatus and the later join-key definition.- Re-running the collector is safe because structurally valid CSV files are revalidated and skipped.
9. On-disk layout
artificial_intelligence_.../
└── POCO_POGO/
└── 2026/
└── document_Document_6a852c3300/
├── page_0001.ris
├── page_0001.csv
├── page_0002.ris
├── page_0002.csv
└── manifest.jsonl
csv_companion_backfill_20260920/
├── queue_summary.json
├── checkpoint.json
├── manifest.jsonl
├── errors.jsonl
├── unresolved.jsonl
└── quarantine/
golden_set_v1/
├── golden_set_master.parquet
├── golden_set_master.csv.gz
├── golden_set_document_eligible.csv.gz
├── validation_sample_100.csv
├── unmatched_ris_records.csv.gz
├── unmatched_csv_records.csv
├── csv_field_conflicts.csv
├── input_csv_manifest.csv
└── SUMMARY.json
The canonical 1950–2026 module view contains relative symbolic links to raw source folders. The backfill resolves each physical source once, preventing the same raw folder from being processed twice when it has more than one catalogue reference.
10. Reproduction procedure
- Use an authorised Policy Commons account and record the exact query URL, modules, filters, page size, sort direction, page number and UTC collection time.
- Install the pinned runtime or record equivalent versions. The worked result used Python 3.14.6, pandas 3.0.5 and pyarrow 25.0.1. The general collectors use Playwright 1.60.0.
- Launch Chrome with a local debugging port and sign in interactively. Do not write cookies or signed addresses to configuration files.
- Freeze RIS, CSV and manifests before integration. Record SHA-256 for every input file.
- Run the structurally bounded CSV collection. During collection require the documented schema, UTF-8 decoding, submitted-item cardinality and atomic
.parthandling. - Place
build_golden_set_v1.pyin the frozen AIGOV5 collection root besideraw_csv/andanalysis_steps5_8_v4/. - Run the builder exactly once, then rerun it to confirm idempotent row counts and hashes for unchanged inputs.
- Require zero duplicate accepted compound keys, zero blank join components, zero temporary files and 100 unique records in the validation sample.
- Review unmatched and conflict files before any later release. Do not delete them merely to improve the match rate.
python -m py_compile collect_partitioned_ris.py \
collect_csv_2026_all_modules.py backfill_csv_companions.py \
build_golden_set_v1.py
python -u collect_csv_2026_all_modules.py
python -u backfill_csv_companions.py
python -u build_golden_set_v1.py
python - <<'PY'
import pandas as pd
d = pd.read_parquet('golden_set_v1/golden_set_master.parquet')
assert len(d) == 11306
assert not d.duplicated(['join_coi','join_artifact_id']).any()
assert d.join_coi.ne('').all() and d.join_artifact_id.ne('').all()
s = pd.read_csv('golden_set_v1/validation_sample_100.csv')
assert len(s) == 100 and not s.record_id.duplicated().any()
PY
Policy Commons holdings and metadata can change. Exact reproduction therefore means executing the frozen method and retaining dated manifests, not assuming that a later query will return identical records.
11. Public code snapshot
- Partitioned RIS collector, SHA-256
ec265bf2ab93456cb181cd4462edf02efd752b80af60c200e6e45f15c7141dc9 - Bounded 2026 CSV companion collector, SHA-256
964b256e617cc7151f93747c48c27fe87ee3093e6720459d085fca27740088d0 - Canonical historical CSV backfill, SHA-256
631457d5e87b3f7e9fe83c685e7a931eb60b2573330a6c59b88ce094677e2eb8 - AIGOV5 identifier-exact Golden Set builder, SHA-256
e5c0f4adc62022387655bfb1583e8e933c79f8aecec6b2637a659e39c31e57d0
The repository contains code and method documentation only. Licensed RIS and CSV payloads, authentication state and signed download addresses are not published.
Citation: Chun, Y. (2026). Paired Policy Commons RIS and CSV collection: authenticated export, identity validation and recovery protocol. Implementation 1.1.