I study how AI systems fail quietly — when a model looks well behaved on aggregate metrics while
its reasoning has been hijacked, its values misaligned with the community it serves, or its actions
unsafe in the physical world. My work builds the measurements that surface those failures.
Current flagship project
Physical AI Risk Assessment for Urban Deployment
MIT Senseable City Lab
Project lead · Frontier AI Lab, KT, with the
MIT Senseable City Lab and the City of Seoul ·
Aug 2026 – present
A twelve-month project that takes the Physical AI Risk Taxonomy out of the document corpus and into
a live city. It builds a risk-assessment framework for how embodied AI and street-level conditions
interact — where delivery robots, autonomous vehicles, and humanoids meet real pavements, crowds,
and weather — and works towards a Korea-specific evaluation standard for urban deployment.
The analytical layer treats the taxonomy as a complex system rather than a tree. The corpus of
1,660 active L4 risks is represented as a weighted semantic network, from which 50 overlapping L3
communities emerge. That representation makes it possible to ask which risks act as hubs, which
semantic pathways connect distant failure modes, and how those structures behave at city scale.
This is where the taxonomy work becomes operational: a structured risk map is only useful if it
survives contact with the street.
Figure 1 | From risk corpus to city street.a, 1,660 active L4 risk
cards across 54 semantic families. b, The corpus embedded as a weighted semantic network,
from which 50 overlapping communities emerge and hub risks become visible. c, Communities
mapped onto the physical locations where embodied AI operates, giving street-level risk exposure
for urban deployment.
AI safety and red-teaming
Red-teaming and defence of retrieval-augmented generation systems
Project lead · KT Frontier AI Lab, with Korea University (Prof. Buru Chang's group) · 2025–2026
Agentic RAG systems resist classic poisoning because they retrieve and reason iteratively,
discarding weakly relevant documents. This programme shows that robustness is thinner than it
looks, and then builds the instrument to measure exactly how thin. It has produced two outputs so
far — an attack, accepted to the EMNLP 2026 Main Conference, and a benchmark —
alongside a registered Korean patent and defences for deployed B2B RAG services.
KidnapRAG — the attack EMNLP 2026
A black-box sequential attack that hijacks an agent's multi-step reasoning chain using three
role-specific poisoned documents: Bait attracts initial retrieval, Chain-Link
redirects query reformulation, and Mal-Ins injects attacker-controlled evidence. It
requires no access to system prompts, reasoning traces, retrievers, or model parameters — only the
ability to publish externally retrievable documents. Accepted to the EMNLP 2026 Main Conference
(15.4% acceptance rate).
Most RAG security benchmarks score only the final response, which says little about where an attack
enters or how it propagates. ARS-Bench measures the entire attack lifecycle — from adversarial
document ingestion to final response generation — across five security-critical dimensions:
document ingestion, retrieval, reasoning, final response, and efficiency. Tracing effects through
the lifecycle enables systematic comparison of attack behaviour, fine-grained diagnosis of
vulnerabilities, and targeted assessment of defences.
Under review, ACL Rolling Review (August 2026)· Benchmark and project resources released with the paper
Governance and risk taxonomy
RAI Risk Taxonomy 2.0
Frontier AI Lab, KT · release v2.18.0-rc
An evidence-linked responsible AI risk taxonomy spanning general-purpose, agentic, and Physical AI.
The release registers 1,711 permanent L4 identifiers, displays 1,660 active bilingual risk cards,
and maps them against 54 semantic L3 families (33 general, 6 agentic, 15 physical) under three L1
domains.
Two design commitments make it auditable. First, L4 identity is permanent and independent of
placement, so a card can be re-classified without losing its provenance. Second, a card with no
exact fit is never forced into the nearest family: it is flagged
needs_taxonomy_decision and enters the RAI-HOLD review queue, where 614 cards
currently await human adjudication. Placement quality is validated through a BGE-M3 sensitivity
analysis run both with and without HOLD cards.
Measuring taxonomy granularity: percolation transitions in semantic space
Project lead · Frontier AI Lab, KT · 2026 · manuscript under submission
How many leaves should a risk taxonomy have? The number is normally an editorial commitment made
before analysis begins, and the merge threshold behind it is set by judgement ("0.8 seems
reasonable"). This project argues that the threshold is a quantity to be measured. Sweeping a
similarity-threshold graph over 1,612 risk cards reveals two percolation-like transitions that
bound the usable range, and an independent criterion adapted from the cohesion species concept in
evolutionary biology, the last point at which within-cluster cohesion still exceeds
between-cluster attraction, selects the same boundary from unrelated statistics.
Merging changes the density of the space, so the boundary moves with it. Iterating merge and
re-derivation generates a granularity flow whose states are the released tiers, and a per-step
null test decides where the trajectory stops rather than an analyst doing so. The same pipeline,
with nothing retuned, reproduces both transitions and a null-valid five-step flow on the public
MIT AI Risk Repository, which indicates that granularity is a property of the inventory rather
than of the embedding scale.
Figure 2 | Why granularity has to be measured.a, One threshold applied
to the whole inventory lets similarity chain: at τ = 0.70, a and d land in the same cluster
through b and c, although their own similarity is 0.50. Any leaf count read off this collapse is
an editorial choice. b, The stepwise flow instead re-derives a crossing boundary at each
step on the contracted graph — τ*1 = 0.8333 separates {a, b} from {c, d}, and
τ*2 = 0.6998 keeps their representatives apart — so the weak 0.50 link never
survives.Figure 3 | The same transitions on an independent inventory. The pipeline
rerun on the MIT AI Risk Repository (database v4, 74 frameworks) with no parameter retuned.
a, Pooled mean pairwise similarity breaks at τ1 = 0.806 ± 0.004,
where similarity chaining sets in; the per-cluster minimum collapses at τ2 =
0.629 ± 0.009. Means over 100 subsample replicates, 95% bands. b, Each
consolidation step lowers the crossing boundary τ*t and the inventory size,
taking 1,810 entries to 727 in five steps. Granularity is a property of the inventory, not of
the embedding scale.
Physical AI risk taxonomy
Project lead · Frontier AI Lab, KT · 2026–present
A public, bilingual evidence map of risks that arise when AI perceives, decides, and acts in the
physical world — embodied AI, humanoids, robots, drones, autonomous vehicles, and cyber-physical
systems. It is built with a hybrid method that pairs a human-defined risk hierarchy with bottom-up
discovery from a large corpus of papers, patents, and policy documents, and released with alignment
tags, audit logs, and a public technical report. The aim is a taxonomy that is auditable rather
than merely descriptive: every card traces back to the evidence that put it there.
The 182 cards serve as the gold-standard layer for the wider taxonomy: 169 exact source
identifiers plus 13 explicit aliases are locked against reclassification, and 360
reference and justification rows with structured 3H and role tags are preserved through every
release. Card quality is checked against anonymous expert annotations collected through a public
survey instrument.
AI Topic Space: policy, science, and technology in the Korean AI knowledge ecosystem
Principal investigator · joint study with STEPI and KDI · 2025–present
An interactive semantic map of AI topics across policy reports, academic papers, and patents, built
on a four-level topic hierarchy. It supports reference-space exploration and relative topic-gap
analysis across five-year periods, making it possible to ask where national policy discourse lags
the research and patent frontier — and where it leads.
Disagreement between a human annotator and an LLM evaluator is usually reported as a single
agreement score, which hides what actually went wrong. I decompose it into two separable
quantities: moral orientation — which values a judge weights — and moral calibration
— how strongly it applies them. In human annotators the two are coupled; in judge LLMs they come
apart. That separation is what makes silent misalignment measurable.
Figure 4 | Two axes of moral disagreement.a, In human annotators,
moral orientation and calibration move together. b, In judge LLMs the two come apart:
a model can weight the right values yet apply them at the wrong strength. That separation is
what makes silent misalignment measurable.
Contextual value alignment through genre-based aesthetic judgement
With Isa Weisha Wang, Institute for Ethics in AI, Technical University of Munich · manuscript in preparation
Where do LLM evaluators diverge from human interpretive communities, and is that divergence
universal or context-dependent? Using film genres as measurable cultural contexts, this framework
distinguishes universal from context-dependent value dimensions and locates human–AI gaps through
a sign-flip test and a subspace-resolved orientation and calibration audit.
A smile blinds the judge: role-blind valence pooling in vision-language models
With Isa Weisha Wang, Institute for Ethics in AI, Technical University of Munich · manuscript in preparation
Vision-language models are increasingly deployed as safety judges, which assumes they read a scene
the way people do, by working out who is doing what to whom. In a fixed scene whose roles are
settled by the narrative, we find they do something cruder: they sum the emotional surface of each
face without regard to who wears it. The victim's smile lowers judged danger and the perpetrator's
smile lowers judged malice, each face blinding its own channel. The effect is significant from 3B
local models up to a 397B frontier model, so scale alone does not supply the missing role binding.
When is AI a cognitive complement? A critical-transition account of dependence
Research programme in development · 2026–
As AI enters human judgement and decision-making, cognition becomes a property of a hybrid
human–AI ensemble rather than of the individual. The question is when AI acts as a
complement that strengthens autonomy and resilience, and when as a substitute
that breeds dependence — and whether that boundary can be measured while aggregate performance
still looks intact.
The hypothesis is that dependence sets in critically rather than gradually.
Model the ensemble as a directed weighted graph whose edges are delegation relations — which
functions are offloaded, how often, how reversibly. Autonomy is then the persistence of
cognitive pathways that bypass the AI subsystem; resilience is the capacity to re-internalise
delegated functions after perturbation. This makes the claim falsifiable: as delegation
density rises, the system should undergo a percolation-like transition at which bypass
pathways disconnect and autonomy collapses abruptly, while aggregate behaviour remains
unchanged.
Knowledge spaces and innovation measurement
The same embedding, clustering, and network machinery that maps a technology space also maps a risk
space. These projects build the large-scale corpora and methods behind that claim, and all are
released so the results can be rebuilt end to end.
DCI Patent Atlas
Data Center Interconnect: corpus, value chain, citation network, and what the capability buys
A patent record for the optical interconnect technology that carries traffic between data
centres, paired with a national panel that asks what such capability converts into. Five
retrieval routes — classification codes, keywords, a curated applicant dictionary,
citation neighbours and targeted recollection — are merged, deduplicated at
application, family and organisation level, and graded rather than filtered, giving a single
master of 1,033,230 applications of which 566,053 are core
(priority years 1980–2025, 157 applicant countries, 86 offices). Every application is
placed on a seven-position value chain, from materials and substrate through
optical components, module assembly and advanced packaging to systems and submarine
operations, plus adjacent technologies, by two independent routes with multiple membership
permitted. The citation layer reconstructs 2,188,730 directed edges over
927,481 applications on a unique-pair basis, correcting an earlier double-counted figure, and
exposes strongly asymmetric national flows: once self-citations are set aside, Japan is the
largest net provider of cited knowledge and the United States, despite being the largest node
on both sides, is the largest net absorber, with China second.
On top of the corpus, a country × field × year panel (35 countries,
2003–2025) links DCI capability to revealed comparative advantage in AI science (2.70M
Web of Science papers, 253 fields) and AI technology (2.33M AI patents, 15 canonical fields,
with the 44,651-application DCI overlap removed from the outcome). DCI capability transfers
directly into adjacent AI patent fields; crossing into AI science requires
an existing scientific base, so the capability alone does not secure comparative advantage in
a new science field; and holding an advantage once gained is explained by DCI proximity on
its own. Advanced packaging is the single value-chain position that matters on both margins.
Research-ready AI patent corpus and collection workflow
A master dataset of 2,330,553 unique AI patent applications with 175 variables,
partitioned by priority-year period and released with processing code, quality metadata, and
checksums. The accompanying collector is a reproducible notebook workflow that automates six
relational query groups against PATSTAT Online within its query-cost, query-length, and
download-size constraints, producing a one-row-per-patent master file.
Earlier work measured how AI capability accumulates across nations and firms — technological
specialisation, AI and green innovation, and design innovation under technology complexity.
See Publications for the journal record.
Policy engagement
Risk taxonomies matter only if they reach the people who write rules. I bring this work into
national and international policy settings.
OECD, Paris — with the KDI delegation
1–2 October 2026
Invited to present with the Korea Development Institute delegation to the OECD, alongside
colleagues from the Bank of Korea, Chung-Ang University and the KDI School. The visit builds
on an earlier exchange with the OECD on a four-dimensional framework for classifying
AI risk — the question of what dimensions a risk classification must carry before it
can support cross-country comparison and policy monitoring.
1 Oct — OECD STI AI and Emerging Digital Technologies Division. A working
meeting on AI risk classification. My talk brings the policy-gap visualisation work to it:
where a country's AI policy agenda falls short of the international reference space, and how
those gaps can be located and measured rather than asserted.
2 Oct — KDI–OECD Expert Roundtable on AI-Driven Innovation, Technology Diffusion and
Productivity, convened with the OECD STI Productivity, Innovation and
Entrepreneurship Division — taking part in the discussion of how AI diffusion translates into
measured productivity, and where risk classification bears on it.
NRC national AI collaborative research programme
Korea, 2026
Policy-Gap Visualisation at the Korean Association for Policy Studies 2026 Summer
Conference, Gwangju (14 Aug 2026), in the STEPI-convened session Inclusiveness and
Policy-Agenda Gaps in the AI Transition: linkage mechanisms, visualisation, and discourse.
Invited talk on the same work at the programme's dissemination forum, The AI Transition of
Industry and Social Institutions, Hotel President, Seoul (7 Oct 2026).
Figure 5 | Where Korean AI policy discourse falls short. Policy L3 topics on
the reference UMAP layout, 2022–2026, graded by relative shortfall
r = normalised Korea activation share ÷ normalised reference activation share. Hollow
circles mark absolute gaps where Korea carries no mass at all; crimson marks strong shortfalls
(r < 0.25). The strongly underrepresented regions are not peripheral — they include
the alignment problem, democratic oversight, surveillance capitalism, gender bias, and the
hard-law/soft-law question, which is the evidence behind the policy-gap talks.
Standing advisory and consulting
AI governance and innovation policy consulting for KDI, STEPI, and Seoul AI Hub; contributor to the
GSMA AI for Impact Taskforce Responsible AI Maturity Roadmap; advisory roles with STEPI's National
Grand Challenges Energy Division and the KOTRA Global Economic and Trade Outlook Report.
Seminars and research notes
Working documents that sit upstream of the projects above: seminar presentations, close
readings of individual papers, and short methodological notes. They are revised as the
underlying work develops. The full archive is at
Research notes.
Interpretable machine learning for preference-coefficient estimation
Seminar
Seminar note · 19 Aug 2026 · presented by
Myongji Cho,
Ph.D. candidate, TEMEP, Seoul National University
How should the heterogeneity of individual preferences be recovered from choice data alone?
The note compares two neural approaches that place their assumptions in different places:
MDN-Logit, which estimates a mixture distribution over attribute coefficients conditioned on
individual characteristics, and MAPL, which abandons attribute-level coefficients and
estimates the distribution of aggregate preference per alternative. Read together they show
that no model removes distributional assumptions — it only moves them — which makes the
location of the assumption an empirical question about the stability of policy indicators
such as the value of travel time.