Back to Education

AI Data Collection

English · 한국어

TIM51001 · Data Collection and Analysis

Collecting artificial intelligence research papers and patents with Codex. These two tutorials end with downloaded records and a collection log. Statistical analysis and full-text acquisition are outside their scope.

Available data and export limits depend on your institutional subscription and account permissions.

1. Web of Science papers · 2. EPO PATSTAT patents

1. Collect AI papers from Web of Science

Establish access and create an account

UNIST students should enter Web of Science through the university library. Institutional access and a personal Web of Science account serve different purposes. Library authentication provides access to subscribed resources; a free personal account alone does not provide a database subscription.

  1. Open the UNIST Library website and sign in with your UNIST credentials. Select English in the header if preferred.
  2. Go to Electronic Resources and then Databases, or open the database list filtered by W.
  3. When working off campus, follow the library's Off Campus Access procedure before opening the database. Check that the status is ON. If access is not active, complete the library's authentication steps first.
  4. Select W in the alphabetical index or search for Web of Science in the database search box. Find Web of Science(SCIE) in the results and open its database link.
  5. Continue through the library's redirect to Web of Science so that institutional access is retained. Check that Web of Science Core Collection and the indexes needed for your study are available. If access is denied, return to the library's off-campus access instructions or contact the library.
UNIST Library database page with W selected, Off Campus Access ON, and the Web of Science SCIE entry near the bottom
W0. Access through UNIST Library. The alphabetical filter is set to W, Off Campus Access is ON in the upper-right header, and Web of Science(SCIE) appears at the bottom. Open the database through this entry before creating or signing in to your personal Web of Science account. UNIST Library database list.

Select Register, enter your email address, name and password, complete the verification step, and activate the account using the email link. Sign in to save searches. Complete passwords and account verification yourself before asking Codex to work in the browser. See Clarivate's registration instructions.

Web of Science registration form with email, password and name fields
W1. Registration. Activate your personal account, then verify institutional access separately. Source.
Derive keywords from previous research

Start with Liu, Shapira and Yue (2021), Tracking developments in artificial intelligence research. Their approach combines benchmark papers, keyword refinement and a subject-category search. Read the methods and keyword tables before asking Codex to draft a query. The short teaching example below is not a reproduction of their validated strategy.

  1. Define your scope. For this exercise, collect AI-related papers published in 2020–2025. Decide whether applications of AI are included and whether conference papers belong in the corpus.
  2. Extract candidate terms from two or three relevant studies. Record each term, its source DOI or URL, table or page, and your reason for including it.
  3. Group synonyms with OR. Use AND only to restrict the AI set to another concept. Avoid the unqualified abbreviation AI, which has unrelated meanings.
  4. Check a small set of known relevant papers and screen 30–50 retrieved titles and abstracts. Record missed papers and false positives; revise the query before the full export.
Concept groupCandidate termsHow to document the choice
Core AIartificial intelligenceUse the seed concept discussed by Liu et al.; record the exact source location.
Learning methodsmachine learning, deep learning, neural network, support vector machineCheck definitions and variants against the studies you read.
Functional scopenatural language processing, computer visionDecide whether the whole function or only learning-based work belongs in your study.
Optional recent extensiongenerative AI, large language modelAdd only with a newer supporting study. Label this as an extension, not a term set taken from a 2021 paper.

Ask Codex to organise the evidence before generating syntax:

Read the methods and search-term tables in the papers I provide.
Create a keyword register with term, synonym, source DOI/URL,
page/table and inclusion rationale. Flag ambiguous terms.
Do not invent citations or attribute new terms to older studies.
Using only the terms I approve, draft a Web of Science Core
Collection query for AI papers published in 2020–2025.
Build and check the search expression

Open Advanced Search → Query Builder. A manageable starter query is shown below. TS searches Topic fields, including title, abstract, author keywords and Keywords Plus; PY sets publication years. Keep the selected citation indexes and document types in your log. AI conference coverage depends on the subscribed indexes. Consult the Query Builder documentation for field tags and search behaviour.

TS=("artificial intelligence" OR "machine learning"
    OR "deep learning" OR "neural network*"
    OR "support vector machine*"
    OR "natural language processing" OR "computer vision")
AND PY=(2020-2025)

For an initial export, change the year range to one year. Expand only after the query and export fields are checked. The example favours clarity over exhaustive coverage; it is a candidate search definition, not a complete census of AI research.

Web of Science Core Collection Query Builder and field-tag controls
W2. Query Builder. Select Core Collection, paste the approved expression, and record the Exact Search setting. Source.
Use Codex to collect and check export batches

Here, browser-assisted collection means asking Codex to operate the search and export controls in an authorised session. Confirm with your library that automation is permitted for your intended volume. Use the provider's export facility where available; a metadata export does not download article PDFs.

From the results page, choose Export and a format such as plain text or tab-delimited text. Select the fullest record fields available, including cited references if required and licensed. Export one small batch first. Confirm that accession identifiers, titles, abstracts and keywords are present. Use the limit shown in your export dialogue rather than assuming a fixed batch size. Clarivate notes that Fast 5000 includes fewer fields. Export documentation.

Web of Science results page showing filters and the Export menu
W3. Results and Export. Use Export to download search results. This wastewater example illustrates the menu; use your own AI query for the exercise. Source.
Use my authorised Web of Science browser session. I have checked
that this collection is permitted by my institution's licence.
Run the approved AI query and record the query, indexes, filters,
search date, result count and export fields in collection_log.json.
Export one small batch first and check that it contains the required
fields. Then export non-overlapping ranges within the displayed
limit, one at a time, to data/wos/raw/ with numbered filenames.
Keep the downloaded files unchanged. Track each completed range,
row count and file checksum so a failed run can resume safely.
Report missing identifiers, duplicate accession IDs and any gap
between the result count and the exported unique records.
Stop on a login challenge, access denial or download restriction.
Do not bypass limits, extract credentials or scrape publisher PDFs.

Finish with unchanged export files, the final search expression, the keyword register and the collection log. Compare unique Web of Science accession IDs with the recorded search total. DOI is a useful secondary identifier but can be absent. Investigate gaps before calling the collection complete.

2. Collect AI patents from EPO PATSTAT

Create access and select the database edition

Start at the EPO PATSTAT page. This tutorial uses PATSTAT Online, the subscription-based SQL interface. The EPO also offers PATSTAT through its free Technology Intelligence Platform (TIP), which has a different workflow.

For a first exercise, follow Free trial, complete the registration details and verification yourself, and follow the access instructions you receive. Select PATSTAT Global and record the edition displayed in the session. Trial access has export restrictions; a registered account alone should not be treated as unrestricted access.

Public EPO PATSTAT Online trial registration page
P1. Trial access. Complete the registration form and email verification. Open registration.
Translate a literature-based definition into patent criteria

Use WIPO Technology Trends 2019: Artificial Intelligence and its search methodology as a starting reference. WIPO distinguishes AI techniques, functional applications and application fields, and combines keywords with classification information. Maintain a patent-specific keyword register rather than copying a paper query unchanged.

In that register, distinguish phrases searched in titles or abstracts from IPC/CPC codes. Verify any code against the classification scheme and database edition. For example, a G06N classification search can supply an additional candidate set, but it is neither equivalent to all AI patents nor necessarily restricted to your definition. Document whether you use keywords alone, keywords OR codes, or keywords AND codes. These produce different populations.

The SQL below is an English title/abstract keyword pilot. It does not implement WIPO's full search strategy. Patent metadata coverage varies, and missing English text can exclude relevant applications. Choose filing year deliberately; it differs from priority year and publication year. Recent filing cohorts may be incomplete because publication and indexing take time.

Write and run a bounded SQL query

Open Search → Expert. In tls201_appln, appln_id identifies the application and docdb_family_id links related applications. Titles are in tls202_appln_title and abstracts in tls203_appln_abstr. Check these fields in the catalogue for your edition.

SELECT TOP 100
    a.appln_id,
    a.docdb_family_id,
    a.appln_auth,
    a.appln_nr,
    a.appln_filing_date,
    a.appln_filing_year,
    t.appln_title,
    b.appln_abstract
FROM tls201_appln AS a
LEFT JOIN tls202_appln_title AS t
    ON t.appln_id = a.appln_id
    AND t.appln_title_lg = 'en'
LEFT JOIN tls203_appln_abstr AS b
    ON b.appln_id = a.appln_id
    AND b.appln_abstract_lg = 'en'
WHERE a.appln_filing_year = 2020
  AND (
    LOWER(t.appln_title) LIKE '%artificial intelligence%'
    OR LOWER(b.appln_abstract) LIKE '%artificial intelligence%'
    OR LOWER(t.appln_title) LIKE '%machine learning%'
    OR LOWER(b.appln_abstract) LIKE '%machine learning%'
    OR LOWER(t.appln_title) LIKE '%deep learning%'
    OR LOWER(b.appln_abstract) LIKE '%deep learning%'
    OR LOWER(t.appln_title) LIKE '%neural network%'
    OR LOWER(b.appln_abstract) LIKE '%neural network%'
  )
ORDER BY a.appln_id;

TOP 100 is a preview, not the full corpus and not a random sample. LIKE '%...%' matches a character sequence; unlike a bibliographic search engine, it does not expand linguistic variants. Repeat approved terms for both title and abstract. A left join retains applications when one text field is missing.

The PATSTAT Online manual specifies that queries must start with SELECT. Its April 2025 revision removes support for CONTAINS and FREETEXT. This example therefore uses LIKE. This is an untested example; run a small trial against your database edition first. If a query exceeds the server cost limit, narrow the date interval or retrieve fields in smaller stages rather than repeating the same expensive query.

PATSTAT manual page 23 showing Expert Search tables, SQL editor, messages and query history
P2. Expert Search. Enter SQL in Query and check Messages for errors after execution. The 2018 Spring interface below may differ from the current layout. Source.
Collect the result table with Codex

Inspect the returned titles and abstracts before extending the pilot. To collect the full chosen interval, remove TOP 100 only after splitting the scope into manageable, non-overlapping date intervals. Save each SQL file with its export. Preserve appln_id and docdb_family_id; multiple applications may belong to the same family. Filing authority is not the applicant's country.

In the Table window, use Download → Prepare download, choose the result table, and retrieve the prepared file from the download manager. Check the limits displayed for your account. Do not confuse a result-table export with a PATSTAT subset export. The manual's download section distinguishes these routes.

PATSTAT manual page 60 showing Prepare download and export formats
P3. Download. Select the result-table route for the rows produced by your SQL. Published manual limits are contextual; the current service and your account determine what is available. Source.
Use my authorised PATSTAT Online browser session. Read the approved
keyword register and the data catalogue for the selected edition.
Generate SELECT-first SQL for the English title/abstract pilot above.
Do not use CONTAINS, FREETEXT or a leading WITH clause.
Run the 100-row pilot and check the fields and sample relevance.
After the pilot, prepare bounded date intervals for the approved
scope and use the supported result-table download controls.
Save SQL and unchanged exports under data/patstat/raw/.
Log edition, query, filing-date interval, query result count,
downloaded row count, completion status and file checksum.
Reconcile the completed intervals and unique appln_id values.
Preserve family IDs without collapsing families at collection time.
On a cost-limit error, narrow the query before retrying.
Stop on authentication challenges or access/export restrictions.
Never infer success solely from clicking the download button.

Finish with the approved patent keyword register, SQL files, downloaded result tables and a collection log. Verify file contents, identifiers and row counts. Keep empty abstracts as missing values and preserve the raw exports before any later cleaning.