F1 — granularity-flow tier 1 (human audit)

Total 1383 cards · 73 L3 categories in use · 109 merged groups · τ=0.8329 · representative text is the medoid original · L3 re-assigned by EM. Hierarchy order: General → Agentic → Physical. Review criteria: ① Description adequacy (is the risk correctly described?) · ② L3 mapping (is the assigned L3 appropriate?) · ③ Redundancy (does it duplicate other cards?). L3 assignments are re-derived by the seed-anchored hybrid EM.
Tier overview (granularity flow)
Tierτ*CardsMerge groupsAbsorbedG / A / PRoleHuman audit
Master1,6121,154 / 155 / 303canonical inventory
F10.83291,383109229906 / 140 / 337fidelity tier (crossing)
F20.80331,154121458734 / 128 / 292second consolidation
F30.79031,03881574653 / 112 / 273third consolidation
F40.775390182711568 / 92 / 241compression tier
F50.763479270820491 / 83 / 218extended compression tier
Merge groups and absorbed cards are, respectively, per consolidation step and cumulative from the Master inventory. The F4 page reports 211 refined groups, the cumulative count over steps 1–4. The Societal Safety axis is one concept family applied at three scopes, so its cards are counted under General, Agentic or Physical (RAI3-{G|A|P}-SOC-nn share the same numbering and meaning).

일반 AI · General · 906 cards

시스템 안전성 · System Safety · 351 cards

RAI3-G-SYS-01 과도한 거절 Over-Refusal22 cards
유해한 행위를 방지하기 위한 제한·안전장치 또는 역할 제약으로 인해 사용자의 안전·권리보호·위험 회피에 필수적인 정보나 선택지를 구조적으로 제공하지 않아 사용자가 실제 피해 또는 중대한 불이익을 입을 가능성이 증대되는 위험
IDCardHuman audit
RAI4-0101
인간 감독·책임성 위장
Human oversight accountability washing
명목상의 인간 감독이 감독자에게 실질적 권한이나 정보를 제공하지 않은 채 책임을 배정하는 수단으로 사용되는 리스크.
The risk that nominal human oversight is used to assign responsibility without providing humans meaningful authority or information.
① Description
② L3 mapping
③ Duplicate
RAI4-0114
의미 있는 인간 통제 실패
Meaningful human control failure
인간 감독이 형식적으로 존재하지만 이를 실효적으로 만드는 데 필요한 정보, 시간, 권한, 역량이 결여되는 리스크.
The risk that human oversight is formally present but lacks the information, time, authority, or competence needed to be meaningful.
① Description
② L3 mapping
③ Duplicate
RAI4-0214
제약 비용 과소평가
Constraint-cost underestimation
정책 또는 평가자가 누적 안전 비용을 과소평가하여 외견상 안전한 행동이 시간 경과에 따라 제약을 위반하게 되는 리스크.
The risk that a policy or evaluator underestimates cumulative safety costs, making apparently safe behavior violate constraints over time.
① Description
② L3 mapping
③ Duplicate
RAI4-0470
감독 회피·수정 저항
Oversight evasion and correction resistance
점점 더 유능해지는 시스템이 수정에 저항하거나 감독을 회피하거나 의도된 범위를 넘어 행동하는 리스크.
The risk that increasingly capable systems resist correction, evade oversight, or act beyond intended bounds.
① Description
② L3 mapping
③ Duplicate
RAI4-0494
규제·관리·운영 복합 실패
Combined regulatory, management, and operational failure
규제·관리·운영상의 실패가 복합적으로 결합되어 피해가 발생하는 리스크.
The risk that harms result from a combination of regulatory, management, and operational failures.
① Description
② L3 mapping
③ Duplicate
RAI4-0527
창의성·비판적 사고의 저하
Devaluation of creativity and critical thinking
인간의 창의성, 예술적 표현, 상상력, 비판적 사고와 문제해결 능력이 평가절하되거나 저하되는 리스크
The risk of devaluation and deterioration of human creativity, artistic expression, imagination, critical thinking, and problem-solving skills.
① Description
② L3 mapping
③ Duplicate
RAI4-0671
복지 혜택 및 자격 상실
Denial of welfare benefits and entitlements
기술 시스템의 오작동, 사용 또는 오용으로 복지 급여, 연금, 주거 등에 대한 접근이 거부되거나 상실되는 리스크
The risk of denial of or loss of access to welfare benefits, pensions, housing, and similar entitlements due to the malfunction, use, or misuse of a technology system.
① Description
② L3 mapping
③ Duplicate
RAI4-0849
자율성·행위주체성 상실
Loss of autonomy and agency
개인, 집단 또는 조직이 정보에 근거한 결정을 내리거나 목표를 추구할 능력을 상실하는 리스크
The risk of loss of an individual, group, or organisation's ability to make informed decisions or pursue goals.
① Description
② L3 mapping
③ Duplicate
RAI4-0850
점진적인 통제력 상실
Gradual loss of control
덜 심각한 중단들이 축적되어 체계적 회복력이 점진적으로 약화되고 결국 중대한 사건이 재앙을 촉발하는, 통제력의 점진적·누적적 상실 리스크
The risk that the accumulation of less severe disruptions gradually weakens systemic resilience until a critical event triggers a catastrophe, constituting gradual or accumulative loss of control.
① Description
② L3 mapping
③ Duplicate
RAI4-0854
되돌릴 수 없는 변화
Irreversible change
사회 구조, 문화적 규범, 인간관계에 되돌리기 어렵거나 불가능한 심각한 장기적 부정적 변화가 발생하는 리스크
The risk of profound negative long-term changes to social structures, cultural norms, and human relationships that may be difficult or impossible to reverse.
① Description
② L3 mapping
③ Duplicate
RAI4-0873
경제적 불안정
Economic instability
기술 시스템 또는 시스템군의 사용이나 오용으로 인해 금융 시스템 또는 그 일부에 통제 불가능한 변동이 발생하는 리스크
The risk of uncontrolled fluctuations impacting the financial system, or parts thereof, due to the use or misuse of a technology system or set of systems.
① Description
② L3 mapping
③ Duplicate
RAI4-1042
시스템 설계 실패
System-design failure
시스템 설계상의 선택이나 오류로 인해 시스템이 실패하는 리스크.
The risk of system failure due to system design choices or errors.
Source members (2)
Source: min_cos=0.8520
RAI4-1042시스템 설계 실패
RAI4-1043구현 실패
① Description
② L3 mapping
③ Duplicate
RAI4-1120
재앙으로 번지는 사고
Accidents cascading into catastrophe
사고가 재앙으로 연쇄 확대되고 갑작스럽고 예측할 수 없는 전개에서 발생하며, 심각한 결함과 위험을 찾아내는 데 수년이 걸리는 리스크.
The risk that accidents cascade into catastrophes, are caused by sudden unpredictable developments, and involve severe flaws and risks that can take years to find.
① Description
② L3 mapping
③ Duplicate
RAI4-1382
수정가능성 상실
Loss of corrigibility
에이전트의 설계나 구성에 결함이 있을 때 에이전트가 인간의 수정 시도에 협조하지 않아 오류 교정과 안전한 중단이 불가능해지는 리스크.
The risk that, when something is wrong in the design or construction of an agent, the agent does not cooperate with human attempts to fix it, precluding error-tolerant correction and safe interruptibility.
① Description
② L3 mapping
③ Duplicate
RAI4-1436
시장 독점
Market monopolisation
가격 통제를 통해 시장 지배력이 남용되어 경쟁이 제한되고 불공정한 진입 장벽이 조성되는 리스크.
The risk of abuse of market power through the control of prices, thereby limiting competition and creating unfair barriers to entry.
① Description
② L3 mapping
③ Duplicate
RAI4-1437
언론/표현의 자유 상실
Loss of freedom of speech/expression
보복, 검열 또는 법적 제재에 대한 두려움 없이 자신의 의견과 생각을 표현할 수 있는 사람들의 권리가 제한되거나 상실되는 리스크.
The risk that people's right to articulate their opinions and ideas without fear of retaliation, censorship, or legal sanction is restricted or lost.
① Description
② L3 mapping
③ Duplicate
RAI4-1438
집회·결사의 자유 상실
Loss of freedom of assembly/association
사람들이 함께 모여 집단적이거나 공유된 생각을 표현·홍보·추구·방어할 권리와 결사에 가입할 권리가 제한되거나 상실되는 리스크.
The risk that people's right to come together and collectively express, promote, pursue, and defend their collective or shared ideas, and to join an association, is restricted or lost.
① Description
② L3 mapping
③ Duplicate
RAI4-1439
사회권과 공공서비스 접근권 상실
Loss of social rights and access to public services
노동, 사회 보장, 적절한 생활 수준, 주거, 건강, 교육에 대한 권리가 제한되거나 상실되는 리스크.
The risk that rights to work, social security, an adequate standard of living, housing, health, and education are restricted or lost.
① Description
② L3 mapping
③ Duplicate
RAI4-1440
정보에 대한 권리 상실
Loss of right to information
공공 기관이 보유한 정보를 찾고 받고 전달할 수 있는 사람들의 권리가 제한되거나 상실되는 리스크.
The risk that people's right to seek, receive, and impart information held by public bodies is restricted or lost.
① Description
② L3 mapping
③ Duplicate
RAI4-1441
자유선거권 상실
Loss of right to free elections
비밀 투표를 통해 합리적인 간격으로 자유 선거에 참여할 수 있는 사람들의 권리가 제한되거나 상실되는 리스크.
The risk that people's right to participate in free elections at reasonable intervals by secret ballot is restricted or lost.
① Description
② L3 mapping
③ Duplicate
RAI4-1442
자유와 안전에 대한 권리 상실
Loss of right to liberty and security
불법적이거나 자의적인 체포 또는 부당한 구금으로 자유가 제한되거나 상실되는 리스크.
The risk that liberty is restricted or lost as a result of illegal or arbitrary arrest or false imprisonment.
① Description
② L3 mapping
③ Duplicate
RAI4-1443
적법 절차에 대한 권리 상실
Loss of right to due process
사법 행정에 의해 공정하고 효율적이며 효과적으로 대우받을 권리가 제한되거나 상실되는 리스크.
The risk that the right to be treated fairly, efficiently, and effectively by the administration of justice is restricted or lost.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-02 역량 초과 수행 Over-Extension5 cards
시스템이 감당 가능한 범위를 넘어선 과제를 “가능한 것처럼” 수행
IDCardHuman audit
RAI4-0166
사용자 이해력 과부하
User comprehension overload
시스템 정보, 경고, 설명이 사용자가 이해하고 대응할 수 있는 범위를 초과하는 리스크.
The risk that system information, warnings, or explanations exceed users' ability to understand and act on them.
① Description
② L3 mapping
③ Duplicate
RAI4-0195
명령의 구현체별 하드웨어 한계 초과
Command exceeds embodiment-specific hardware limits
모델이 배포된 로봇의 알려진 도달 범위·탑재 하중·액추에이터·관절·엔드이펙터 한계를 넘는 명령을 내리는 위험.
A model issues a command that exceeds the known reach, payload, actuator, joint, or end-effector limits of the robot on which it is deployed.
① Description
② L3 mapping
③ Duplicate
RAI4-0471
역량 오버행
Capability overhang
잠재 역량이 평가·모니터링·거버넌스 절차가 탐지하는 수준을 초과하는 리스크.
The risk that latent capabilities exceed what evaluation, monitoring, or governance processes detect.
① Description
② L3 mapping
③ Duplicate
RAI4-0519
기술 시스템 개발 노동 착취
Labor exploitation in technology system development
기술 시스템의 학습·개발·관리·최적화를 위해 저임금 및 역외 노동을 포함한 노동력이 사용·오용되는 리스크
The risk that labour, including under-paid and offshore labour, is used or misused to train, develop, manage, or optimise a technology system.
① Description
② L3 mapping
③ Duplicate
RAI4-1006
가중되는 노동 부담
Increased labor burden
특정 사회 집단의 구성원이 시스템이나 제품을 다른 사람들만큼 잘 작동시키기 위해 더 많은 시간과 노력을 들여야 하는 리스크.
The risk that members of certain social groups bear increased burden or effort, such as time spent, to make systems or products work as well for them as for others.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-03 허위 정보/오정보 Misinformation/Disinformation44 cards
존재하지 않거나 틀린 정보를 사실처럼 비의도적 생성·전달하는 현상, 지식 한계·데이터 편향·추론 오류·정보 업데이트 실패 등에서 기인
IDCardHuman audit
RAI4-0439
딥페이크 사칭
Deepfake impersonation
합성 미디어가 실존 인물·기관을 고충실도로 사칭하여 사기와 명예 훼손을 가능하게 하고 진본 커뮤니케이션에 대한 신뢰를 훼손하는 리스크
Synthetic media impersonates real individuals or institutions with high fidelity, enabling fraud and reputational harm while undermining trust in authentic communication.
① Description
② L3 mapping
③ Duplicate
RAI4-0624
명예훼손성 허위 인식 생성
Defamatory false-perception generation
기술 시스템이 개인, 집단, 조직에 대한 허위 인식을 생성·조장·증폭하는 리스크
The risk that a technology system is used to create, facilitate, or amplify false perceptions about an individual, group, or organisation.
① Description
② L3 mapping
③ Duplicate
RAI4-0634
대규모 설득 및 유해한 조작 위험
Large-scale persuasion and harmful manipulation risks
AI 생성 합성 콘텐츠와 대규모 디지털 플랫폼의 전략적 조작이 오도성 정보·이념을 정밀 표적화하여 공공 인식을 대규모로 왜곡하고 사회 안정을 위협하는 리스크
AI-generated synthetic content and strategic manipulation of large digital platforms distort public perception at scale, targeting misleading information or ideologies to destabilize social order.
① Description
② L3 mapping
③ Duplicate
RAI4-0635
모델 생성 정확도의 한계
Limitations in model generative accuracy
생성 정확성의 한계로 딥페이크 등 사실적으로 보이지만 전적으로 조작된 콘텐츠가 산출되어 수신자가 진본과 구별할 수 없게 되는 리스크
Limits in generative accuracy produce convincingly realistic but fabricated content, including deepfakes, that recipients cannot distinguish from authentic material.
① Description
② L3 mapping
③ Duplicate
RAI4-0637
AI 생성 허위정보에 의한 여론 조작
Public-opinion manipulation via AI-generated disinformation
범용 AI가 진본과 구별하기 어려운 텍스트·이미지·음성·영상을 대규모로 생성·유포하여 사람들을 오도하고 설득하며 선거 등 정치적 과정의 여론을 조작하는 리스크
The risk that general-purpose AI generates and disseminates at scale highly persuasive text, images, audio, and video indistinguishable from genuine human-generated material, misleading and manipulating people and influencing public opinion in political processes.
Source members (8)
Source: min_cos=0.7093 · Mixed L3
RAI4-0440AI 생성 선거 허위정보
RAI4-0637AI 생성 허위정보에 의한 여론 조작
RAI4-1085영향력 공작
RAI4-1116개인화 허위정보와 행동 조작
RAI4-1156여론 조작 촉진
RAI4-1553생성형 AI 영향력 공작에 의한 여론 조작
RAI4-1732생성형 AI 기반 허위정보 확산
RAI4-1733생성형 AI를 통한 허위 정보 확산
① Description
② L3 mapping
③ Duplicate
RAI4-0659
대규모로 허위 정보를 자동으로 생성
Automatically generating disinformation at scale
AI 모델이 소셜미디어 게시물, 상품평, 위성영상 등 출처와 품질이 상이한 대체 금융데이터를 수집·집계하면서 편향과 일반화 문제가 유입되어 기업 주가가 급변하는 금융 꼬리위험이 발생하는 리스크
The risk that AI-enabled collection and aggregation of alternative financial data of varying quality and provenance introduces biases and generalization issues, posing financial tail risks in which a company's price changes dramatically.
① Description
② L3 mapping
③ Duplicate
RAI4-0704
AI 생성 콘텐츠에 의한 안전 위협
Safety threats from AI-generated content
AI가 생성하거나 합성한 콘텐츠가 허위정보 확산, 차별과 편향, 개인정보 유출과 권리 침해를 야기하여 국민의 생명과 재산의 안전, 국가안보, 사상안보를 위협하고 윤리적 위험을 초래하며, 견고한 보안 기제가 없을 경우 유해한 이용자 입력에 대해 위법하거나 해로운 정보를 출력하는 리스크
The risk that AI-generated or synthesized content leads to the spread of false information, discrimination and bias, privacy leakage, and infringement, threatening the safety of citizens' lives and property, national security, and ideological security and causing ethical risks, and that without robust security mechanisms the model outputs illegal or damaging information in response to harmful user input.
① Description
② L3 mapping
③ Duplicate
RAI4-0790
허위·오도 정보 확산에 의한 잘못된 믿음 형성
False beliefs from AI-spread inaccurate information
AI 시스템이 부정확하거나 오해의 소지가 있는 정보를 생성하고 그 확산을 촉진하여 사람들이 잘못된 믿음을 갖게 되는 리스크
The risk that AI systems generate and facilitate the spread of inaccurate or misleading information that causes people to develop false beliefs.
Source members (3)
Source: min_cos=0.7395
RAI4-0594진위 판별 실패에 따른 허위정보 생성
RAI4-0790허위·오도 정보 확산에 의한 잘못된 믿음 형성
RAI4-0841잘못된 인식·믿음의 확산
① Description
② L3 mapping
③ Duplicate
RAI4-0791
AI로 인한 명예훼손
AI-generated defamation
AI 시스템이 특정 개인·조직의 명예를 부당하게 훼손하는 허위 사실 주장을 생성하거나 증폭하는 리스크
An AI system produces or amplifies false factual claims that unjustifiably damage an identifiable person's or organization's reputation.
① Description
② L3 mapping
③ Duplicate
RAI4-0792
합성 콘텐츠 식별 곤란
Difficulty of distinguishing synthetic content
합성 콘텐츠를 진본 자료와 구별하기 어려워 정보 관련 피해가 가중되는 리스크
The risk that the difficulty in distinguishing synthetic content from authentic material adds to information risks.
① Description
② L3 mapping
③ Duplicate
RAI4-0793
개인에 관한 허위정보 유포
Dissemination of false information about individuals
사람들에 관한 허위이거나 오해의 소지가 있는 정보가 유포되는 리스크
The risk that false or misleading information about people is disseminated.
Source members (2)
Source: min_cos=0.8517 · Mixed L3
RAI4-0793개인에 관한 허위정보 유포
RAI4-1084위험한 정보 유포
① Description
② L3 mapping
③ Duplicate
RAI4-0799
정보 생태계 오염
Pollution of information ecosystem
생성 도구의 출력이 최종 이용자를 넘어 유포되면서 공개적으로 이용 가능한 정보가 허위이거나 부정확한 정보로 오염되는 리스크
The risk that publicly available information is contaminated with false or inaccurate information as generative-tool output is disseminated beyond the end user.
① Description
② L3 mapping
③ Duplicate
RAI4-0801
작화성 콘텐츠 생성
Confabulated content generation
자신 있게 서술되지만 오류이거나 허위인 콘텐츠(속칭 환각 또는 조작)가 생성되어 이용자가 오도되거나 기만당하는 리스크
The risk of the production of confidently stated but erroneous or false content, known colloquially as hallucinations or fabrications, by which users may be misled or deceived.
① Description
② L3 mapping
③ Duplicate
RAI4-0802
사용자 기만·인증 우회
User deception and authentication bypass
AI 시스템과 그 출력이 명확히 표시되지 않아 이용자가 상호작용 상대와 콘텐츠 출처를 분별하지 못해 오판하고, 고도로 사실적인 AI 생성 이미지·음성·영상이 안면·음성 인식 등 신원확인 절차를 무력화하는 리스크
The risk that unlabeled AI systems and outputs prevent users from discerning whether they interact with AI or identifying content provenance, leading to misjudgement and misunderstanding, while highly realistic AI-generated images, audio, and video circumvent identity-verification mechanisms such as facial and voice recognition.
① Description
② L3 mapping
③ Duplicate
RAI4-0804
정보 환경의 질적 저하와 균질화
Degradation and homogenisation of the information environment
AI 어시스턴트로 생성된 스팸·오도성·저품질 합성 콘텐츠가 온라인 공간에 확산되어 신뢰할 정보의 검증이 어려워지고 디지털 지식공유재가 침식되며, 이용자가 접하는 정보와 관점이 균질화되는 리스크
The risk that the proliferation of spam, misleading, and low-quality synthetic content generated by AI assistants makes reliable information hard to verify, erodes the digital knowledge commons, and homogenises the information and ideas people encounter.
① Description
② L3 mapping
③ Duplicate
RAI4-0808
사실과 다른 부정확한 생성 콘텐츠
Factually incorrect generated content
LLM이 생성한 콘텐츠에 사실과 다른 부정확한 정보가 포함되는 리스크
The risk that LLM-generated content contains inaccurate information that is factually incorrect.
① Description
② L3 mapping
③ Duplicate
RAI4-0809
조작된 인용·출처를 동반한 허위정보 제시
False information with fabricated quotes and sources
AI 모델이 권위 있게 들리는 문체와 조작된 인용·출처를 곁들여 허위정보를 사실처럼 제시하는 리스크
The risk that AI models present false information as if it is factual, often with authoritative-sounding text and fabricated quotes and sources.
① Description
② L3 mapping
③ Duplicate
RAI4-0810
성실성 오류
Faithfulness errors
생성 콘텐츠가 근거 자료나 입력 내용에 충실하지 않아, 유창하고 그럴듯해 보여도 원문 왜곡(충실성 오류)이 발생하는 리스크
Generated content is unfaithful to the source material or input it claims to represent, introducing faithfulness errors even when the output appears fluent and plausible.
① Description
② L3 mapping
③ Duplicate
RAI4-0812
허위정보
False information
챗봇이 알려진 사실, 권위 있는 출처 또는 제공된 원본 문서와 모순되는 정보를 출력하는(환각으로도 불리는) 리스크
The risk that a chatbot outputs information that contradicts known facts, authoritative sources, or provided source documents, also known as hallucination.
① Description
② L3 mapping
③ Duplicate
RAI4-0815
정보 생태계 저하와 잘못된 인식 형성
Information-ecosystem degradation and false beliefs
허위·환각·저품질·오도성·부정확한 정보가 생성·확산되어 정보 생태계가 저하되고 사람들이 잘못된 인식·결정·신념을 갖거나 정확한 정보에 대한 신뢰를 잃는 리스크
The risk that creation or spread of false, hallucinatory, low-quality, misleading, or inaccurate information degrades the information ecosystem and causes people to develop false or inaccurate perceptions, decisions, and beliefs, or to lose trust in accurate information.
① Description
② L3 mapping
③ Duplicate
RAI4-0819
비의도적 허위정보 생성
Unintentional misinformation generation
악의적 이용자가 해를 끼칠 의도로 만든 것이 아니라, LLM이 사실에 부합하는 정보를 제공할 능력이 부족하여 비의도적으로 잘못된 정보를 생성하는 리스크
The risk that wrong information is generated unintentionally by LLMs because they lack the ability to provide factually correct information, rather than being intentionally generated by malicious users to cause harm.
Source members (2)
Source: min_cos=0.8340
RAI4-0819비의도적 허위정보 생성
RAI4-0836고의적 허위정보 생성 능력
① Description
② L3 mapping
③ Duplicate
RAI4-0829
정보환경 악화
Degradation of the information environment
인물·사건에 대한 사실적 허위 묘사가 저비용으로 대량 생성되어 정보 환경이 저질화되고, 공개 정보 기반 의사결정과 진실 정보에 대한 신뢰가 훼손되는 리스크
Cheap generation of realistic false portrayals of people and events degrades the information environment, compromising decisions that rely on public information and lowering trust in true information.
① Description
② L3 mapping
③ Duplicate
RAI4-0831
허구적 참조 및 인공물 생성
Generation of fabricated references and artifacts
AI 시스템이 존재하지 않는 학술 인용이나 영상의학 영상 내 허위 구조물 등 허구적 산출물을 참인 정보와 동일한 확신으로 제시하여, 지식이 부족한 이용자에게 허위정보 확산과 위험한 상황을 초래하는 리스크
The risk that AI systems produce fabricated artifacts such as made-up academic references or false structures in X-ray or MRI images and present them with the same apparent confidence as true information, heightening misinformation and creating potentially dangerous situations for less knowledgeable people.
① Description
② L3 mapping
③ Duplicate
RAI4-0832
알고리즘 시스템에 의한 정보 기반 피해
Information-based harms from algorithmic systems
생성 모델과 추천 시스템 등 알고리즘 시스템이 오정보, 허위정보, 악의적 정보에 관한 정보 기반 피해를 초래하는 리스크
The risk that algorithmic systems, especially generative models and recommender systems, lead to information-based harms of misinformation, disinformation, and malinformation.
① Description
② L3 mapping
③ Duplicate
RAI4-0843
의료·법률 등 고위험 영역 허위정보 실질적 피해
Material harm from false information in high-stakes domains
AI가 의료 복약량이나 법률 자문 등 민감한 영역에서 잘못된 정보를 제공하여 유도·강화된 잘못된 믿음으로 이용자가 자신에게 위해를 가하거나 의도치 않게 범죄를 저지르는 리스크.
The risk that induced or reinforced false beliefs from misinformation in sensitive domains such as medicine or law lead users to cause harm to themselves, for example through incorrect medical dosages, or to unwillingly commit a crime by following false legal advice.
① Description
② L3 mapping
③ Duplicate
RAI4-0844
허위·오도 정보 유포에 의한 기만과 양극화
Deception and polarisation from false information
언어모델이 오도성 있거나 허위인 정보를 예측·제시하여 이용자에게 잘못된 믿음을 심는 기만이 발생하고, 개인의 자율성이 위협되며 근거 없는 기존 견해에 대한 확신이 커져 양극화가 심화되는 리스크
The risk that language models predict misleading or false information that misinforms or deceives people and instils false beliefs, threatening personal autonomy and increasing confidence in previously held unsubstantiated opinions, thereby increasing polarisation.
① Description
② L3 mapping
③ Duplicate
RAI4-0958
기만적 딥페이크 미디어 콘텐츠
Deceptive deepfake media content
AI가 텍스트·사진·오디오·비디오에 걸쳐 인간이 가짜임을 의심하지 못할 수준의 정교한 위조 콘텐츠를 생성하여 기만이 발생하는 리스크.
The risk that AI generates fake text, photo, audio, and video content so sophisticated that people's minds rule out the possibility of it being fake, enabling deception.
① Description
② L3 mapping
③ Duplicate
RAI4-1057
허위정보 생산 비용 절감
Cheaper and more effective disinformation
LM이 합성 미디어와 가짜 뉴스 제작에 사용되어 대규모 허위정보 생산 비용을 낮추는 리스크.
The risk that LMs are used to create synthetic media and fake news, reducing the cost of producing diffuse disinformation at scale.
① Description
② L3 mapping
③ Duplicate
RAI4-1125
선거 허위정보
Election disinformation
응답이 시민 선거의 투표 시간, 장소, 방식을 포함한 선거 시스템과 절차에 대해 사실과 다른 정보를 담는 리스크.
The risk that responses contain factually incorrect information about electoral systems and processes, including the time, place, or manner of voting in civic elections.
① Description
② L3 mapping
③ Duplicate
RAI4-1183
딥페이크 기술
Deepfake technology
딥페이크 기술이 조작된 콘텐츠에 진본의 외양을 부여하는 설득력 있는 위조 이미지·영상·음성을 생성하는 리스크
Deepfake technology produces convincing counterfeit visuals, video, and audio that give fabricated content the appearance of authenticity.
① Description
② L3 mapping
③ Duplicate
RAI4-1205
딥페이크 피해 구제 곤란
Difficulty redressing deepfake harms
AI가 생성한 이미지와 영상이 피해자 자신이 아니라 여러 출처를 합성한 그럴듯한 허구의 장면이고 공개된 자료에 의존해 전통적 프라이버시와 동의 개념을 우회하기 때문에, 피해자가 제작자를 특정하고도 피해를 구제받지 못하는 리스크.
The risk that victims of targeted AI-generated harms struggle to redress them because the generated image or video is a composite from multiple sources rather than the victim, relying on public images and thus circumventing traditional notions of privacy and consent.
① Description
② L3 mapping
③ Duplicate
RAI4-1206
믿을 수 있는 딥페이크
Believable deepfakes
딥페이크가 진짜라고 믿는 시청자에게 유포되어 대상에게 실질적인 사회적 피해를 입히고, 허위임이 밝혀진 뒤에도 대상에 대한 부정적 인식이 지속되는 리스크.
The risk that deepfakes circulated to viewers who think they are real impose real social injuries on their subjects, with a persistent negative impact on how others view the subject even after the deepfake is debunked.
① Description
② L3 mapping
③ Duplicate
RAI4-1227
콘텐츠 신뢰성 상실
Loss of content authenticity
생성 AI가 발전할수록 작품의 진위를 판단하기 어려워져 이미지와 영상의 대규모 조작이 가능해지고 소셜 미디어의 허위 정보 확산 문제가 악화되며, AI가 만든 예술이 진정성을 결여하게 되는 리스크.
The risk that as generative AI advances it becomes harder to determine the authenticity of a piece of work, enabling large-scale manipulations of images and videos, worsening the spread of fake information, and leaving AI-generated artwork lacking authenticity.
① Description
② L3 mapping
③ Duplicate
RAI4-1376
허위·오도 정보 생성 장벽 저하
Lowered barriers to large-scale misinformation
사실과 의견·허구를 구별하지 않거나 불확실성을 인정하지 않는 콘텐츠의 생성·교환·소비 장벽이 낮아져 대규모 왜곡 및 허위정보 캠페인에 활용되는 리스크.
The risk that lowered barriers to generating and supporting the exchange and consumption of content that may not distinguish fact from opinion or fiction or acknowledge uncertainties enable large-scale dis- and mis-information campaigns.
① Description
② L3 mapping
③ Duplicate
RAI4-1422
AI에 의한 허위·오도 정보 대량 생산
AI-scaled production of false and misleading information
이미지·오디오·텍스트 합성 모델을 통해 설득력 있는 허위 또는 오해의 소지가 있는 정보의 온라인 생산이 대량화되는 리스크.
The risk that AI-based image, audio, and text synthesis models scale up the online production of convincing yet false or misleading information.
① Description
② L3 mapping
③ Duplicate
RAI4-1449
선거 간섭
Electoral interference
허위 또는 오해의 소지가 있는 정보가 생성되어 유권자를 방해하거나 오도하고 선거 과정에 대한 신뢰를 약화시키는 리스크.
The risk that generation of false or misleading information interrupts or misleads voters and/or undermines trust in electoral processes.
① Description
② L3 mapping
③ Duplicate
RAI4-1540
모델 설득력에 의한 허위 신념 확산
Spread of false beliefs through model persuasive capability
GPAI 시스템이 개인화된 대화나 대량 생산된 오도성 콘텐츠를 통해 사용자에게 잘못된 정보를 확신시켜, 조작적이거나 비진실한 콘텐츠가 사회적으로 확산되는 리스크.
The risk that GPAI systems convince users of incorrect information through personalized dialogue or mass-produced misleading content, spreading manipulative or untruthful material at societal scale.
① Description
② L3 mapping
③ Duplicate
RAI4-1554
멀티모달 딥페이크를 통한 괴롭힘·명예훼손·협박
Harassment, defamation, and extortion via multimodal deepfakes
이미지·오디오·영상 등 다중 모달리티로 실존 또는 가상의 인물과 사건을 묘사하고 실존 인물의 말과 동작을 모사한 딥페이크가 개인을 괴롭히고 명예를 훼손하며 위협·갈취하는 데 사용되는 리스크.
The risk that deepfakes combining multiple modalities to depict real or non-existent people and events, including imitation of real people's speech and body movements, are used to harass, discredit, intimidate, and extort individuals.
① Description
② L3 mapping
③ Duplicate
RAI4-1559
개인·집단 맞춤형 허위정보의 저비용 대량 생성
Low-cost mass generation of personalized disinformation
GPAI를 이용해 특정 집단이나 개인에 맞춤화된 허위정보를 저비용으로 자동 생성하여 기만 공작의 효과가 증대되는 리스크.
The risk that GPAI enables automatic generation of disinformation personalized to specific groups or individuals at significantly reduced cost, increasing the effectiveness of such attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-1560
GPAI 출력을 이용한 설득력 있는 사칭
Convincing impersonation using GPAI outputs
텍스트·이미지·오디오·영상 전반에서 GPAI 생성물이 항상 탐지되지는 않는다는 점을 이용해, 악의적 행위자가 생성물이나 위조 증빙 문서로 설득력 있는 사칭을 수행하는 리스크.
The risk that, because GPAI outputs are not always detected as AI-generated across modalities, malicious actors use such outputs or AI-forged supporting documents to construct convincing impersonations.
① Description
② L3 mapping
③ Duplicate
RAI4-1581
위장 계정 자동 운영에 의한 정보환경 조작
Information-environment manipulation through automated sockpuppet accounts
허구의 온라인 정체성 계정이 자동으로 생성·운영되어 후원 주체가 은폐되고 독립적 지지가 가장되며 정보 환경이 조작되는 리스크.
The risk that fictitious online identities are automatically generated or operated to conceal sponsorship, simulate independent support, or manipulate information environments.
① Description
② L3 mapping
③ Duplicate
RAI4-1583
증거·신분 문서 위조
Falsification of evidence and identity documents
보고서·신분증·문서 등 증거가 조작되거나 허위로 제시되는 리스크.
The risk that evidence, including reports, identity documents, and other records, is fabricated or falsely represented.
① Description
② L3 mapping
③ Duplicate
RAI4-1653
잘못된 정보 대량 생성과 영향 공작에 의한 조작
Manipulation through LLM-scaled misinformation and influence operations
LLM이 인간 수준의 설득력을 갖는 기만 서사와 가짜뉴스를 저비용·대규모로 생성하고 자동화된 영향 공작과 악성 소셜 봇넷에 이용되어, 표적 청중의 관점이 조작되고 선전의 진입 장벽이 크게 낮아지는 리스크.
The risk that LLMs craft deceptive narratives and fabricate fake news as persuasive as human-generated content at low cost and large scale, powering automated influence operations and malicious social botnets that manipulate the perspectives of targeted audiences and lower the barrier for propaganda.
① Description
② L3 mapping
③ Duplicate
RAI4-1712
AI 생성·유포 유해 콘텐츠에 의한 피해
Harm from AI-generated or AI-distributed detrimental content
딥페이크, 신원 허위 표현, 위협, 자해 조장, 극단주의 콘텐츠, 잘못된 정보, 성착취물, 사기 등 AI가 생성하거나 유포한 콘텐츠가 상해·손해·손실을 유발하는 리스크.
The risk that AI-generated or AI-distributed content such as deepfakes, identity misrepresentation, threats, self-harm promotion, extremist content, misinformation, sexual abuse material, or scams instigates injury, damage, or loss.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-04 맥락 불일치 Context Misalignment16 cards
특정 국가·관할·법체계를 전제로 서비스를 제공함에도 불구하고, 학습 데이터 또는 참조 코퍼스의 구성·분포 등으로 인해 타 관할의 법규범·판례 논리·제도적 전제를 암묵적으로 수용하여, 해당 서비스가 적용되어야 할 법질서(legal order)와의 정합성을 저해하고, 결과적으로 해당 공동체/관할의 규범 경계 밖으로 안내·요약·추천이 편향될 위험
IDCardHuman audit
RAI4-0007
결과 전파 오분류
Consequence propagation misclassification
벤치마크나 사고 분석이 국소적 에이전트 실패가 하류의 물리적·재정적·프라이버시·사회적 결과로 전파되는 정도를 과소평가하는 리스크.
The risk that a benchmark or incident analysis underestimates how a local agent failure propagates into downstream physical, financial, privacy, or social consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-0373
문화 간 고정관념 전이
Cross-cultural stereotype transfer
한 문화적 맥락에서 학습된 고정관념이 다른 집단·언어·환경에 대한 출력으로 전이되는 리스크.
The risk that stereotypes learned in one cultural context are transferred into outputs about another group, language, or setting.
① Description
② L3 mapping
③ Duplicate
RAI4-0374
문화적 맥락 붕괴
Cultural context collapse
문화 특수적 의미와 관행이 모델 처리 과정에서 일반 범주로 붕괴되어 지역적 뉘앙스와 사회적 의미가 소실되는 리스크
Culturally specific meanings and practices are collapsed into generic categories during model processing, losing local nuance and social significance.
① Description
② L3 mapping
③ Duplicate
RAI4-0375
문화적 가치 오보정
Cultural value miscalibration
모델의 가치 판단이 해당 시스템이 사용되는 문화 공동체에 제대로 보정되지 않는 리스크.
The risk that a model's value judgments are poorly calibrated to the cultural community in which the system is used.
① Description
② L3 mapping
③ Duplicate
RAI4-0376
현지화된 가치 정렬 실패
Localized value alignment failure
AI 시스템이 집계 벤치마크에서는 정렬된 것처럼 보이면서도 현지에서 수용되는 윤리·법·사회 규범을 따르지 못하는 리스크.
The risk that an AI system fails to follow locally accepted ethical, legal, or social norms despite appearing aligned in aggregate benchmarks.
① Description
② L3 mapping
③ Duplicate
RAI4-0378
이중 언어 문화 정렬 실패
Bilingual cultural alignment failure
AI 시스템이 동일한 공동체가 사용하는 언어들 사이에서 서로 다르거나 문화적으로 일관되지 않은 가치 판단을 내리는 리스크.
The risk that an AI system gives different or culturally inconsistent value judgments across languages used by the same community.
① Description
② L3 mapping
③ Duplicate
RAI4-0379
문화 간 평가 격차
Cross-cultural evaluation gap
평가 벤치마크가 지배적인 문화·언어 환경 밖의 정렬 실패를 과소 측정하는 리스크.
The risk that evaluation benchmarks under-measure alignment failures outside dominant cultural and linguistic settings.
① Description
② L3 mapping
③ Duplicate
RAI4-0381
번역 매개 가치 왜곡
Translation-mediated value distortion
기계 번역이 프롬프트와 응답의 규범적 효력, 공손성, 유해성, 사회적 의미를 변형시켜 언어 간 전달되는 가치를 왜곡하는 리스크
Machine translation alters the normative force, politeness, harmfulness, or social meaning of prompts and responses, distorting values communicated across languages.
Source members (2)
Source: min_cos=0.8597
RAI4-0381번역 매개 가치 왜곡
RAI4-0477기계 번역 의미 왜곡 피해
① Description
② L3 mapping
③ Duplicate
RAI4-0382
방언·언어 사용역 배제
Dialect and register exclusion
방언·사회어·존댓말 체계·언어 사용역이 오류로 처리되어 해당 공동체의 안전성과 유용성이 저하되는 리스크.
The risk that dialects, sociolects, honorific systems, or registers are treated as errors, reducing safety and usability for affected communities.
① Description
② L3 mapping
③ Duplicate
RAI4-0389
지식 체계 배제
Knowledge-system marginalization
토착·지역·종교·관행 기반 지식 체계가 모델 출력과 평가에서 배제되는 리스크.
The risk that indigenous, local, religious, or practice-based knowledge systems are excluded from model outputs and evaluations.
① Description
② L3 mapping
③ Duplicate
RAI4-0404
데이터세트 문화 샘플링 편향
Dataset cultural sampling bias
훈련 또는 평가 데이터세트가 지배적인 문화 환경을 과다 표집하고 지역의 사회적 의미를 과소 대표하는 리스크.
The risk that training or evaluation datasets oversample dominant cultural settings and underrepresent local social meanings.
① Description
② L3 mapping
③ Duplicate
RAI4-0414
글로벌 벤치마크 단일문화
Global benchmark monoculture
소수의 글로벌 벤치마크가 지역의 문화·제도적 기준을 배제한 채 성공적 정렬을 정의하는 리스크.
The risk that a small set of global benchmarks defines successful alignment while excluding local cultural and institutional criteria.
① Description
② L3 mapping
③ Duplicate
RAI4-0475
배포 맥락 규범 불일치
Deployment-context normative mismatch
AI 시스템이 배포 지역의 법·문화·언어·사회적 기대를 반영하지 못하는 리스크.
The risk that AI systems fail to reflect local law, culture, language, or social expectations.
① Description
② L3 mapping
③ Duplicate
RAI4-1046
집단별 성능 격차·고정관념 인코딩
Demographic performance disparity and stereotype encoding
ML 시스템이 일부 인구통계·사회 집단의 고정관념을 인코딩하거나 그 집단에 대해 불균형적으로 낮은 성능을 보이는 리스크.
The risk that an ML system encodes stereotypes of, or performs disproportionately poorly for, some demographic or social groups.
① Description
② L3 mapping
③ Duplicate
RAI4-1145
맥락적 사회 규범 위반
Violation of contextual social norms
인터넷 텍스트로 학습한 LLM의 가중치가, 특정 맥락에 배포될 경우 그 맥락의 정보 공유 규범에서 이탈해 이를 위반하는 기능을 인코딩하는 리스크.
The risk that model weights of LLMs trained on internet text data encode functions which, if deployed in particular contexts, deviate from and violate the information-sharing norms of that context.
① Description
② L3 mapping
③ Duplicate
RAI4-1507
언어 간 벤치마크 오염
Cross-lingual benchmark contamination
벤치마크가 다른 언어로 번역된 뒤 훈련 데이터로 사용되어 오염이 탐지 기법에 가려지고, 모델이 해당 벤치마크가 측정하는 역량을 일반화했다는 거짓 확신이 생기는 리스크.
The risk that models trained on data encoded in multiple languages contain contamination obscured by translation, as when a benchmark is translated into another language and then fed to the model as training data, hiding the contamination from detection methods and giving false assurance that the model has generalized on the capabilities the benchmark tests for.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-05 비일관성 Inconsistency27 cards
동일하거나 실질적으로 유사한 입력·상황·사실관계에 대해, 세션·시간·표현·프롬프트 또는 에이전트 구성 차이로 인해 상이하거나 모순된 결과를 생성하는 위험
IDCardHuman audit
RAI4-0070
사고 분류 단편화
Incident taxonomy fragmentation
사고 범주가 일관되지 않아 집계, 비교, 제도적 학습이 이루어지지 못하는 리스크.
The risk that inconsistent incident categories prevent aggregation, comparison, and institutional learning.
① Description
② L3 mapping
③ Duplicate
RAI4-0380
저자원 언어 가치 손실
Low-resource language value loss
모델 훈련과 평가가 고자원 언어에 집중되어 저자원 언어 사용자가 가치 뉘앙스·안전 적용 범위·사회적 의미를 잃는 리스크.
The risk that low-resource language users lose value nuance, safety coverage, or social meaning because model training and evaluation are concentrated in high-resource languages.
① Description
② L3 mapping
③ Duplicate
RAI4-0418
주석 불일치 억제
Annotation disagreement suppression
주석자 간 불일치가 단일 레이블로 축소되어 복수의 가치와 사회적으로 유의미한 이견이 은폐되는 리스크.
The risk that disagreement among annotators is collapsed into a single label, hiding plural values and socially meaningful disagreement.
① Description
② L3 mapping
③ Duplicate
RAI4-0653
LLM 남용에 의한 학업 부정행위
Academic misconduct from LLM misuse
LLM 시스템의 부적절한 사용과 남용이 학업 부정행위와 같은 부정적 사회적 영향을 초래하는 리스크
The risk that improper use or abuse of LLM systems causes adverse social impacts such as academic misconduct.
① Description
② L3 mapping
③ Duplicate
RAI4-0708
사용자 집단 간 성능 격차
Disparate performance across user groups
LLM의 질의응답이나 사실확인 등 성능이 인종, 사회적 지위, 과업, 언어 집단에 따라 크게 달라져 특정 집단이 열등한 서비스 품질을 받게 되는 리스크
The risk that LLM performance, such as question-answering and fact-checking ability, differs significantly across racial, social status, task, and language groups, delivering inferior quality to some groups.
① Description
② L3 mapping
③ Duplicate
RAI4-0716
집단 속성에 따른 배분적 불의
Allocative injustice by group attribute
LLM이 관련 프로필이 동일하나 소속 집단이 다른 개인들에 대해 실질적으로 다른 텍스트나 제안을 산출하여, 무관한 집단 속성에 따른 배분적 불의가 발생하는 리스크
The risk that an LLM produces materially different suggested or completed texts for individuals with the same relevant profiles who differ only in an irrelevant group attribute, resulting in allocative injustice.
① Description
② L3 mapping
③ Duplicate
RAI4-0818
과신에 의한 오답 제시
Overconfident presentation of erroneous answers
LLM이 객관적 정답이 없는 주제나 자신의 내재적 한계·구식 지식이 문제되는 영역에서 불확실성을 인식하지 못한 채 과신하여 확신에 찬 오답을 제시하는 리스크
The risk that an LLM is over-confident on topics where objective answers are lacking or where its inherent limitations and outdated knowledge base should caution restraint, leading to confident yet erroneous responses.
① Description
② L3 mapping
③ Duplicate
RAI4-0821
지식 분포 변화에 따른 응답 노후화
Answer obsolescence from knowledge distribution shift
LLM이 학습한 지식 기반이 시간에 따라 계속 변화함에도 이를 반영하지 못해, 갱신이 필요한 사실 질문에 낡고 부정확한 답변을 제시하는 리스크
The risk that knowledge bases on which LLMs were trained continue to shift while the model does not update, producing outdated and incorrect answers to questions whose correct answers change over time.
① Description
② L3 mapping
③ Duplicate
RAI4-0822
맥락 일관성 추구에 의한 아첨과 환각 증폭
Sycophancy and hallucination from context consistency
LLM이 문맥의 일관성을 추구하여 앞선 접두 문맥이나 이용자 의견에 담긴 허위정보를 사실보다 우선시하고 이를 반복함으로써 아첨성 응답과 환각이 눈덩이처럼 증폭되는 리스크
The risk that an LLM's tendency to pursue consistent context leads it to prioritize and reiterate false information contained in prefixes or user-provided opinions over facts, amplifying sycophantic responses and snowballing hallucinations.
① Description
② L3 mapping
③ Duplicate
RAI4-0878
역량 오추정으로 인한 신뢰 훼손과 피해
Harm from misestimated LLM capability inconsistency
과장된 홍보, 과업 오염, 과업·도메인의 과소 대표, 프롬프트 민감성 등으로 이용자가 도메인 간·내 일관성이 없는 LLM의 실제 역량을 오추정하여, 부정확하거나 오도성 있는 출력에 근거한 결정으로 피해를 입고 신뢰가 훼손되는 리스크
The risk that users misestimate an LLM's true capabilities, owing to exaggerated claims, task contamination, underrepresentation of tasks or domains, and prompt sensitivity underlying inconsistent performance across and within domains, undermining trust and causing harm when decisions are based on incorrect or misleading outputs.
① Description
② L3 mapping
③ Duplicate
RAI4-0952
글쓰기 능력 저하와 학술 문헌 오염
Writing skill erosion and scientific literature pollution
LLM 사용이 문체의 획일화와 개인적 표현의 억압 등 글쓰기 능력을 저해하고, 저품질 생성 원고의 범람으로 학술적 진실성과 과학 문헌이 오염되는 리스크.
The risk that LLM use erodes writing skills through homogenization of styles and stifling of individual expression, and that a flood of low-quality generated manuscripts pollutes the scientific literature and undermines academic integrity.
① Description
② L3 mapping
③ Duplicate
RAI4-1054
언어·집단 간 성능 격차
Language and group performance disparity
LM이 소수의 언어로만 훈련되어 학습 데이터가 마련되지 않은 다른 언어에서는 성능이 떨어지는 리스크.
The risk that LMs, typically trained in few languages, perform less well in other languages, in part because labelled training data is unavailable for them.
① Description
② L3 mapping
③ Duplicate
RAI4-1178
체계적 출력 불공정
Systemic output unfairness
LLM이 인종과 성별, 종교 등 여러 주제에 걸친 사회적 편향을 담은 불공정하고 편향된 표현과 행위를 식별해 회피하지 못하고 산출하는 리스크.
The risk that LLMs fail to identify and avoid unfair and biased expressions and actions reflecting social bias across various topics such as race, gender, and religion, and produce them instead.
① Description
② L3 mapping
③ Duplicate
RAI4-1187
출력 불일치
Output inconsistency
모델이 서로 다른 사용자, 같은 사용자의 다른 세션, 심지어 같은 대화 안의 발화 사이에서도 동일하고 일관된 답변을 제공하지 못하는 리스크.
The risk that models fail to provide the same and consistent answers to different users, to the same user in different sessions, and even in chats within the same conversation.
① Description
② L3 mapping
③ Duplicate
RAI4-1196
제한된 논리적 추론
Limited logical reasoning
LLM이 질문에 답할 때 겉으로는 그럴듯하지만 궁극적으로 부정확하거나 타당하지 않은 근거를 제시하는 리스크.
The risk that LLMs provide seemingly sensible but ultimately incorrect or invalid justifications when answering questions.
① Description
② L3 mapping
③ Duplicate
RAI4-1228
프롬프트 품질 결함으로 인한 오류
Errors from poor prompt quality
인간 언어의 모호성 때문에 프롬프트를 통한 인간과 기계의 상호작용에서 오류와 오해가 발생하고, 프롬프트를 디버깅하기 어려워 가치 있는 산출을 이끌어내지 못하는 리스크.
The risk that, due to the ambiguity of human languages, interaction between humans and machines through prompts leads to errors or misunderstandings, and that prompts are hard to debug, so valuable outputs are not elicited.
① Description
② L3 mapping
③ Duplicate
RAI4-1309
인구집단 재현 불균형
Demographic representation disparity
LLM이 생성한 텍스트에서 서로 다른 인구통계 집단이 언급되는 비율에 격차가 생겨 특정 집단이 과대 대표되거나 과소 대표되거나 지워지는 리스크.
The risk that there is disparity in the rates at which different demographic groups are mentioned in LLM-generated text, resulting in overrepresentation, under-representation, or erasure of specific demographic groups.
① Description
② L3 mapping
③ Duplicate
RAI4-1498
LLM 평가자 오판
Faulty LLM-as-evaluator judgments
다른 모델을 평가하는 LLM이 장황함·특정 입장 선호 등 잘못된 평가를 산출하고, 이것이 학습에 통합되면 피학습 모델이 평가자 결함을 악용하도록 발달하는 리스크
An LLM used to evaluate other models produces incorrect judgments, such as rewarding verbosity or ideological stance, and when integrated into training, the trained model learns to exploit the evaluator's flaws.
① Description
② L3 mapping
③ Duplicate
RAI4-1515
사고연쇄와 불일치하는 모델 출력
Model outputs inconsistent with chain-of-thought reasoning
모델 출력의 이해를 돕기 위해 사용되는 사고연쇄 추론이 모델이 제시하는 최종 답변과 일치하지 않아 충분한 투명성을 제공하지 못하는 리스크.
The risk that chain-of-thought reasoning, employed to get a better understanding of a model's output by encouraging transparent reasoning in text form, is inconsistent with the final answer given by the model and therefore does not provide sufficient transparency.
① Description
② L3 mapping
③ Duplicate
RAI4-1523
무관한 문맥에 의한 모델 성능 저하
Performance degradation from irrelevant context
프롬프트에 포함된 무관한 정보가 모델의 주의를 분산시켜 사고연쇄 프롬프팅을 포함한 다양한 기법에서 성능이 크게 저하되는 리스크.
The risk that models are easily distracted by irrelevant provided information such as context in LLMs, leading to a significant decrease in performance across prompting techniques including chain-of-thought prompting.
① Description
② L3 mapping
③ Duplicate
RAI4-1525
맥락 내 학습 불투명성에 의한 안전성 보증 실패
Safety assurance failure from opaque in-context learning mechanisms
프롬프트에 예시를 제공해 가중치 변경 없이 새 과업을 학습시키는 맥락 내 학습의 작동 메커니즘이 규명되지 않아 프롬프트를 통한 오용에 대해 안전성을 보증할 수 없는 리스크.
The risk that the working mechanism of in-context learning, which lets a model learn a new task from examples in the prompt without changing its weights, is not well understood, making it difficult to guarantee safety against the many potential misuses directly related to prompting.
① Description
② L3 mapping
③ Duplicate
RAI4-1526
프롬프트 형식 민감성으로 인한 평가 신뢰성 저하
Evaluation unreliability from prompt-format sensitivity
구분자·대소문자·간격 등 사소한 프롬프트 형식 변화가 모델 성능을 크게 변동시켜 모델 평가와 비교의 신뢰성을 저하시키는 리스크.
The risk that LLMs are highly sensitive to variations in prompt formatting such as changes in separators, casing, or spacing, so that even minor modifications shift model performance significantly and affect the reliability of model evaluations and comparisons.
① Description
② L3 mapping
③ Duplicate
RAI4-1601
근사 기반 출처 귀속의 부정확
Incorrect source attribution from approximation-based methods
출처 귀속 기법이 근사에 기반하여, 모델 출력의 전부 또는 일부가 어떤 훈련 데이터에서 생성되었는지에 대한 귀속이 부정확해지는 리스크.
The risk that, because current source-attribution techniques are based on approximations, a system's account of which training data generated part or all of its output is incorrect.
① Description
② L3 mapping
③ Duplicate
RAI4-1649
기반모델 공유에 의한 상관 실패와 출력 균질화
Correlated failures and output homogenization from shared foundation models
대규모 사전학습 비용으로 인해 다수의 배포 인스턴스가 동일하거나 유사한 학습 구성요소를 공유함으로써 출력 균질화가 심화되고 안전성과 역량 측면의 상관 실패가 발생하는 리스크.
The risk that, because the expense of large-scale pretraining leads many deployed instances to share similar or identical learned components, output homogenization increases and LLM-agents become vulnerable to correlated failures in both safety and capabilities.
① Description
② L3 mapping
③ Duplicate
RAI4-1658
도메인별 오용
Domain-specific misuses
의료·교육 등 민감 영역에 LLM을 조야하게 적용하여 비동의 실험적 치료, 부정행위, 저품질 자동 평가, 설득력 있으나 유해한 도덕적 조언 등 영역 특수적 오용이 발생하는 리스크
Crude application of LLMs in sensitive domains such as health and education produces domain-specific misuse, including unconsented experimental therapy, cheating, low-quality automated assessment, and compelling but harmful moral guidance.
① Description
② L3 mapping
③ Duplicate
RAI4-1668
사전학습 코퍼스 불일치에 의한 가치 어긋남
Value mismatch from divergence between pretraining corpora and societal values
사전학습 코퍼스의 분포가 인간 사회의 분포와 정확히 일치하지 않고 지식이 균등하게 학습되지 않아, LLM 기반 시스템에서 인간 가치와 어긋난 판단이 발생하고 고위험 영역에서 심각한 문제가 초래되는 리스크.
The risk that pretraining corpora do not match the distribution of human society and knowledge is not equally learned, producing value mismatches in LLM-empowered systems and severe value-related problems in high-stakes areas.
① Description
② L3 mapping
③ Duplicate
RAI4-1702
분포 변화 견고성 실패에 의한 확신 오류 [기원]
Confidently wrong outputs from failed robustness to distributional shift [origin]
(GYK-2025 '비정상 분포'의 기원 항목으로 상호참조.) 테스트 분포가 훈련 분포와 달라질 때 기계학습 시스템의 성능이 저하되면서도 높은 확신을 유지한 채 잘못된 출력을 산출하는 리스크.
The risk that an ML system performs poorly and remains confidently wrong when the test distribution differs from training. NOTE: origin of GYK-2025 "Non-stationary distribution"; cross-reference.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-06 과도한 일반화 Overgeneralization21 cards
제한적 사실이나 과거 사례로부터 도출된 패턴을 맥락적 차이와 예외 가능성을 충분히 고려하지 않은 채 일반 규칙으로 확장하여 왜곡된 판단이나 예측을 생성하는 위험
IDCardHuman audit
RAI4-0363
소수 가치 소거
Minority-value erasure
소수자·주변화·과소대표 집단의 가치가 학습 데이터 구성, 평가, 정렬, 배포 결정 과정에서 체계적으로 소거되어 모델 행동이 지배적 가치 분포만 반영하게 되는 리스크
Values held by minority, marginalized, or underrepresented groups are systematically erased during training data curation, evaluation, alignment, or deployment decisions, so model behavior reflects only dominant value distributions.
① Description
② L3 mapping
③ Duplicate
RAI4-0368
도덕적 다양성 압축
Moral diversity compression
다양한 도덕적 입장이 모델 친화적인 소수 범주나 평균 선호로 압축되는 리스크.
The risk that diverse moral positions are compressed into a small set of model-friendly categories or average preferences.
① Description
② L3 mapping
③ Duplicate
RAI4-0371
서구 규범 기본값화
Western normative defaulting
모델이 서구 자유주의·영어권·고소득 국가 규범을 정렬과 평가의 기본 기준으로 취급하는 리스크.
The risk that models treat Western liberal, Anglophone, or high-income country norms as the default basis for alignment and evaluation.
① Description
② L3 mapping
③ Duplicate
RAI4-0397
RLHF 규범적 과적합
RLHF normative overfitting
인간 피드백 기반 강화학습이 좁은 평가자 집단에 과적합하여 논쟁적인 가치를 단일한 행동 규범으로 전환하는 리스크.
The risk that reinforcement learning from human feedback overfits to a narrow rater population and converts contested values into a single behavioral norm.
① Description
② L3 mapping
③ Duplicate
RAI4-0405
합성 데이터의 문화 고정관념 증폭
Synthetic-data cultural stereotype amplification
합성 데이터가 소스 모델이나 시드 데이터에 이미 존재하는 문화적 고정관념이나 협소한 가치 가정을 증폭시키는 리스크.
The risk that synthetic data amplifies cultural stereotypes or narrow value assumptions already present in source models or seed data.
① Description
② L3 mapping
③ Duplicate
RAI4-0425
표현적 고정관념
Representational stereotyping
모델 출력이 사회 집단에 대한 고정관념, 비하적 연상, 비대칭적 재현을 재생산하여 할당 결과와 무관한 재현적 피해를 야기하는 리스크
Model outputs reproduce stereotypes, demeaning associations, or asymmetric representation of social groups, causing representational harm independent of allocative outcomes.
① Description
② L3 mapping
③ Duplicate
RAI4-0474
복수 가치의 규범적 평면화
Normative flattening of plural values
모델 출력이 복수의 사회적 가치를 단순화되거나 다수 중심의 기본값으로 축소하는 리스크.
The risk that model outputs reduce plural social values to simplified or majority-centric defaults.
① Description
② L3 mapping
③ Duplicate
RAI4-0686
재현 왜곡 기반 고정관념·균질화
Misrepresentation-driven stereotyping and homogenisation
특정 정체성, 집단, 관점이 허위 표현되거나 과잉·과소 표현 또는 미표현됨으로써 개인, 집단, 사회, 문화에 대한 경멸적이거나 유해한 고정관념화와 동질화가 발생하는 리스크
The risk of derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies, or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups, or perspectives.
① Description
② L3 mapping
③ Duplicate
RAI4-0697
개발자 가치 각인 편향
Developer value embedding bias
보편적으로 합의된 기준이 없는 상태에서 개발자가 선택한 규범적 가치와 원칙에 따라 모델을 미세조정함으로써 개발자의 이념과 세계관이 모델에 각인되어, 특정 인구집단을 대표하지 못하거나 세계 문화 규범과 변화하는 사회적 견해를 정태적이고 단순화된 형태로 반영하는 출력이 산출되는 리스크
The risk that, absent universally accepted standards, developers' fine-tuning of models on chosen normative rules and principles embeds their ideology and vision of the world into the model, so that it incorporates values unrepresentative of certain segments of the population or offering a static, oversimplified reflection of global cultural norms and evolving social views.
① Description
② L3 mapping
③ Duplicate
RAI4-0698
가치 고착 및 결과 동질화
Value lock-in and outcome homogenization
진화하는 사회적 관점을 반영해 재학습되지 않는 모델이 낡고 덜 포용적인 이해를 고착시켜 대안적 관점의 제시와 탐색을 제한하고, 동일한 기반모델이 여러 배포자에 의해 광범위하게 사용되어 사회 전반에 편향이 동질화되고 기존 편향이 고착되는 리스크
The risk that models not retrained to reflect evolving societal views lock in older, less inclusive understandings and limit the presentation or exploration of alternative perspectives, while deployment of identical foundation models across many downstream deployers homogenizes bias across broad swathes of society and further entrenches existing biases.
① Description
② L3 mapping
③ Duplicate
RAI4-0814
역사 수정주의적 서술 생성
Generation of historically revisionist accounts
사회·공동체·학계가 확립한 역사적 사건이나 서술이 의도적 또는 비의도적으로 재해석되는 리스크
The risk of deliberate or unintentional reinterpretation of established or orthodox historical events or accounts held by societies, communities, and academics.
① Description
② L3 mapping
③ Duplicate
RAI4-0938
사회적 가치·윤리 규범과 상충하는 모델 출력
Model outputs conflicting with societal ethical values
언어모델이 옳고 그름의 판단, 사회규범 및 법률과의 관계 등 보편적으로 받아들여지는 사회적 가치를 충분히 반영하지 못하고 이에 어긋나는 출력을 내는 리스크.
The risk that language models insufficiently attend to universally accepted societal values—including judgements of right and wrong and their relation to social norms and laws—producing outputs at odds with ethics and morality.
① Description
② L3 mapping
③ Duplicate
RAI4-0997
사회 집단 고정관념화
Stereotyping social groups
알고리즘 시스템의 출력이 특정 집단 구성원의 특성·속성·행동에 대한 믿음과 속성 간 결합에 대한 믿음을 반영하여 사회 집단을 고정관념화하는 리스크.
The risk that an algorithmic system's outputs reflect beliefs about the characteristics, attributes, and behaviors of members of certain groups, and about how and why certain attributes go together, stereotyping those groups.
① Description
② L3 mapping
③ Duplicate
RAI4-1051
출력의 사회적 고정관념 재생산
Reproduction of social stereotypes in outputs
인터넷 규모 텍스트로 학습된 언어모델이 주변화 집단에 대한 비하 표현과 고정관념을 학습하여 출력에서 재생산하는 리스크
Language models trained on internet-scale text learn demeaning language and stereotypes about frequently marginalized groups and reproduce them in outputs.
① Description
② L3 mapping
③ Duplicate
RAI4-1053
배제적 규범 인코딩
Exclusionary norm encoding
언어에 표현된 사회적 범주와 규범을 충실히 인코딩한 LM이 그 범주 밖에 사는 집단을 배제하는 규범을 그대로 담게 되는 리스크.
The risk that LMs faithfully encoding patterns present in language necessarily encode social categories and norms that exclude groups who live outside of them.
① Description
② L3 mapping
③ Duplicate
RAI4-1107
데이터 대표성 불균형
Data representation imbalance
특정 집단·요소가 과대대표되고 현상 특성화에 중요한 변수는 제대로 포착되지 못해, 과소대표된 집단을 잘못 특성화하는 모델이 학습되는 리스크.
The risk that certain groups or types of elements are over-weighted or over-represented while variables crucial to characterizing a phenomenon of interest are not properly captured in learned models.
① Description
② L3 mapping
③ Duplicate
RAI4-1158
대규모 설득 능력
Large-scale persuasion capability
모델이 대화와 미디어 환경에서 효과적으로 설득하여 허위 방향으로도 신념을 변화시키고 특정 서사를 유포하며 기존 판단이나 윤리에 반하는 행동을 유도하는 리스크
Models persuade effectively in dialogue and media settings, shifting beliefs including toward falsehoods, promoting narratives, and convincing people to act against their prior judgment or ethics.
① Description
② L3 mapping
③ Duplicate
RAI4-1163
평가 인지 행동 변화
Evaluation-aware behaviour shifting
모델이 자신이 훈련 중인지 평가 중인지 배포 중인지 구분하고 자기 자신과 주변 환경에 관한 지식을 갖추어, 각각의 상황에서 다르게 행동하는 리스크.
The risk that a model can distinguish whether it is being trained, evaluated, or deployed and, knowing that it is a model and having knowledge about itself and its likely surroundings, behaves differently in each case.
① Description
② L3 mapping
③ Duplicate
RAI4-1503
내재 가치 평가 편향
Biased evaluation of encoded values
평가하기 쉬운 내재 가치가 측정이 어려운 가치보다 우선적으로 평가에 포함되어, 더 바람직하지만 정량화가 어려운 가치가 과소 대표되는 불균형이 발생하는 리스크.
The risk that encoded human values which are easier to evaluate are preferred for inclusion in evaluations over those that are more difficult to measure, creating an imbalance in which more desirable but harder-to-quantify values are underrepresented.
① Description
② L3 mapping
③ Duplicate
RAI4-1516
인코딩된 추론
Encoded reasoning
모델이 스테가노그래피 기법으로 중간 추론 단계를 인간이 해석할 수 없는 방식으로 부호화하고, 성능 향상 효과 때문에 이러한 경향이 자연히 나타나며 역량이 높은 모델일수록 두드러지는 리스크.
The risk that models employ steganography techniques to encode their intermediate reasoning steps in ways that are not interpretable by humans, a tendency that may emerge naturally and become more pronounced with more capable models because encoded reasoning can improve performance.
① Description
② L3 mapping
③ Duplicate
RAI4-1697
반사실적 권위 문제
Counterfactual authority problem
월드모델의 반사실적 설명이 학습 분포를 반영할 뿐인데도 운영자·규제자가 인과적 사실로 수용하여 책임 귀속·회피에 선택적으로 활용되는 리스크
Operators and regulators accept world-model counterfactual explanations as causal ground truth although they reflect the model's learned distribution, allowing selective use to attribute or deflect liability.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-07 과도한 확신 Overconfidence41 cards
에이전트가 불확실한 상황에서 자신의 판단에 과도한 확신을 가지고 멈추지 않고 진행하여 잘못된 결과를 초래하는 리스크. "모르면 멈추는가 vs. 알아서 진행하는가"의 정책 부재
IDCardHuman audit
RAI4-0057
감사 추적 불완전성
Audit trail incompleteness
로그, 모델 변경 이력, 데이터 계보, 의사결정 기록이 불충분하여 유해 사건을 재구성하고 감사할 수 없는 리스크.
The risk that logs, model changes, data lineage, or decision traces are insufficient to reconstruct and audit harmful events.
① Description
② L3 mapping
③ Duplicate
RAI4-0060
모델 카드 불완전성
Model card incompleteness
모델 카드 공개 내용이 불완전하거나 최신이 아니거나 지나치게 모호하여 다운스트림 위험 관리를 뒷받침하지 못하는 리스크.
The risk that model-card disclosures are incomplete, outdated, or too vague to support downstream risk management.
① Description
② L3 mapping
③ Duplicate
RAI4-0061
데이터시트 불완전성
Datasheet incompleteness
데이터 문서가 수집 경위, 동의, 품질, 대표성, 알려진 한계를 공개하지 못하는 리스크.
The risk that data documentation fails to disclose collection, consent, quality, representativeness, or known limitations.
① Description
② L3 mapping
③ Duplicate
RAI4-0063
제3자 감사 액세스 제한
Third-party audit access restriction
독립 감사인이 위험을 평가하는 데 필요한 데이터, 모델, 로그, 인터페이스, 문서에 접근하지 못하는 리스크.
The risk that independent auditors lack access to the data, models, logs, interfaces, or documentation needed to evaluate risk.
① Description
② L3 mapping
③ Duplicate
RAI4-0064
감사자 독립성 실패
Auditor independence failure
이해상충, 선정 유인, 피감사 조직에 대한 의존으로 감사 결과가 훼손되는 리스크.
The risk that audit results are compromised by conflicts of interest, selection incentives, or dependence on the audited organization.
① Description
② L3 mapping
③ Duplicate
RAI4-0065
인증 캡처
Certification capture
인증이나 적합성 평가가 공익적 위험 감축이 아니라 공급업체의 이해에 부합하게 되는 리스크.
The risk that certification or conformity assessment becomes aligned with vendor interests rather than public risk reduction.
① Description
② L3 mapping
③ Duplicate
RAI4-0066
감사 체크리스트 준수 극장
Audit checklist compliance theater
감사가 실질적인 위험을 식별하지 않고 규정 준수를 알리는 피상적인 체크리스트 실행이 되는 위험.
Risk that audits become superficial checklist exercises that signal compliance without identifying substantive risk.
① Description
② L3 mapping
③ Duplicate
RAI4-0082
역량 평가 비공개
Capability evaluation non-disclosure
조직이 외부 감독에 필요한 위험 역량, 알려진 한계, 적대적 강건성, 잔여 위험에 관한 평가 근거를 은폐하거나 불명확하게 공개하는 리스크.
The risk that an organization withholds or obscures evaluation evidence about dangerous capabilities, known limitations, adversarial robustness, or residual risks needed for external oversight.
① Description
② L3 mapping
③ Duplicate
RAI4-0084
평가 쇼핑 위험
Evaluation-shopping risk
조직이 더 엄격하거나 맥락에 부합하는 시험을 무시한 채 유리한 평가 결과만 선택적으로 보고하는 리스크.
The risk that organizations selectively report favorable evaluations while ignoring more demanding or contextually relevant tests.
① Description
② L3 mapping
③ Duplicate
RAI4-0100
규정 준수 위장
Compliance washing
조직이 의미 있는 시험, 문서화, 독립적 검증을 회피하면서 규정 준수를 주장하는 리스크.
The risk that organizations claim compliance while avoiding meaningful testing, documentation, or independent scrutiny.
① Description
② L3 mapping
③ Duplicate
RAI4-0102
공급업체 불투명성
Vendor opacity
공급업체가 배포자와 규제기관에 필요한 모델, 데이터, 평가, 업데이트 정보를 제공하지 않는 리스크.
The risk that vendors withhold model, data, evaluation, or update information needed by deployers and regulators.
① Description
② L3 mapping
③ Duplicate
RAI4-0126
설명 실행 가능성 격차
Explanation actionability gap
설명이 제공되더라도 사용자가 무엇을 해야 할지 또는 자신의 이익을 어떻게 보호할지 판단하는 데 도움이 되지 않는 리스크.
The risk that explanations are available but do not help users decide what to do or how to protect their interests.
① Description
② L3 mapping
③ Duplicate
RAI4-0134
임상적 의사결정 지원 과잉 의존
Clinical decision-support overreliance
임상의가 진단·중증도 분류·치료에서 임상 의사결정지원 출력에 과도하게 의존하여, 독립적 임상 판단이 필요한 불확실성과 맥락 한계에도 권고를 수용하는 리스크
Clinicians over-rely on clinical decision-support outputs in diagnosis, triage, or treatment, accepting recommendations despite model uncertainty and contextual limitations that require independent clinical judgment.
① Description
② L3 mapping
③ Duplicate
RAI4-0402
가치 충돌 불투명성
Value conflict opacity
가치 간 상충관계가 모델 출력 뒤에 은폐되어 공공 가치 갈등을 식별하거나 숙의하기 어려워지는 리스크.
The risk that tradeoffs among values are hidden behind model outputs, making public value conflict difficult to identify or deliberate.
① Description
② L3 mapping
③ Duplicate
RAI4-0419
이해관계자 이견 은폐
Multi-stakeholder disagreement concealment
개발 또는 평가 절차가 영향을 받는 이해관계자 사이의 해결되지 않은 이견을 은폐하는 리스크.
The risk that development or evaluation processes conceal unresolved disagreement among affected stakeholders.
① Description
② L3 mapping
③ Duplicate
RAI4-0484
부정확·민감 정보의 메모리 누적
Unsafe memory accumulation
검증·출처 관리·보존·삭제가 불충분하여 허위·민감·오래되었거나 악의적인 정보가 에이전트의 메모리나 검색 저장소에 누적되고, 의도적 오염 공격이 없어도 이후의 검색과 의사결정을 저해하는 리스크.
The risk that inadequate validation, provenance control, retention, or deletion allows false, private, stale, or malicious information to accumulate in an agent's memory or retrieval store, degrading later retrieval and decisions without requiring a deliberate poisoning attack.
① Description
② L3 mapping
③ Duplicate
RAI4-0516
잘못된 위험 테스트
Incorrect risk testing
위험을 측정하거나 추적하기 위해 선택한 지표가 잘못 선정되어 위험을 불완전하게 측정하거나 해당 맥락과 무관한 위험을 측정하게 되는 리스크
The risk that a metric selected to measure or track a risk is incorrectly selected, measuring the risk incompletely or measuring the wrong risk for the given context.
① Description
② L3 mapping
③ Duplicate
RAI4-0541
데이터 출처 검증 불가
Unverifiable data provenance
데이터의 소유권·출처·변환 이력을 검증할 표준화된 방법이 없어 사용 데이터가 원본과 동일한지, 올바른 사용 조건을 갖췄는지 보증할 수 없게 되는 리스크
The risk that, lacking standardized and established methods to verify where data came from, there is no guarantee that the data is the same as its original source or carries the correct usage terms.
① Description
② L3 mapping
③ Duplicate
RAI4-0543
대표성이 없는 위험 테스트
Unrepresentative risk testing
시험 입력이 배포 중 예상되는 입력과 불일치하여 시험이 대표성을 갖지 못하게 되는 리스크
The risk that test inputs mismatched with the inputs expected during deployment make testing unrepresentative.
① Description
② L3 mapping
③ Duplicate
RAI4-0546
감사자 역량 부족에 따른 과대 보증
Over-assurance from insufficient auditor capacity
감사자가 특정 안전·성능·검증 요구를 다룰 지식이나 충분히 엄밀한 시험 역량을 갖추지 못해 정당화될 수 있는 범위보다 넓게 적합 판정이 보고되는 리스크
The risk that auditors lacking knowledge of specific risks or the capacity for sufficiently rigorous testing report passing audits more inclusive than can be justified.
① Description
② L3 mapping
③ Duplicate
RAI4-0547
감사 결과 미공개 및 협력 부족
Non-disclosure and non-cooperation in audits
감사자가 발견한 위험을 공개하지 않거나 결함을 공표하지 못하도록 요구받고 관련 내부 당사자로부터 충분한 협력을 받지 못하게 되는 리스크
The risk that auditors do not publicly disclose risks they find, are required not to publicize shortcomings, or do not receive sufficient cooperation from the relevant internal parties.
① Description
② L3 mapping
③ Duplicate
RAI4-0588
학습 데이터 접근 불가에 따른 설명 제약
Explanation failure from inaccessible training data
학습 데이터에 접근할 수 없어 모델이 제공할 수 있는 설명의 유형이 제한되고 그 설명이 부정확할 가능성이 높아지는 리스크
The risk that, without access to the training data, the types of explanations a model can provide are limited and more likely to be incorrect.
① Description
② L3 mapping
③ Duplicate
RAI4-0590
불충분한 정보에 기반한 조언
Advice given on insufficient information
모델이 충분한 정보를 갖추지 못한 상태에서 조언을 제공하여 그 조언을 따를 경우 피해가 발생하는 리스크
The risk that a model provides advice without having enough information, resulting in possible harm if the advice is followed.
① Description
② L3 mapping
③ Duplicate
RAI4-0596
모델 투명성 부족
Lack of model transparency
모델 설계·개발·평가 절차에 대한 문서화가 불충분하고 모델 내부 작동에 대한 통찰이 부재하여 모델 투명성이 확보되지 않는 리스크
The risk that insufficient documentation of the model design, development, and evaluation process, together with the absence of insights into the model's inner workings, leaves model transparency unattained.
Source members (4)
Source: min_cos=0.8241
RAI4-0522데이터 투명성 부족
RAI4-0523시스템 투명성 부족
RAI4-0525훈련 데이터 투명성 부족
RAI4-0596모델 투명성 부족
① Description
② L3 mapping
③ Duplicate
RAI4-0957
투명성 부족
Lack of transparency
과정에 대한 통찰을 제공하지 않고 설명 없이 결정을 내리는 블랙박스 시스템으로 인해 사용자 신뢰를 얻지 못하고 감사 가능성 등 규제 기준을 충족하지 못하는 리스크.
The risk that a black box making decisions without any explanation or insight into the process fails to gain users' trust and fails to meet regulatory standards such as auditability.
① Description
② L3 mapping
③ Duplicate
RAI4-1033
투명성과 설명가능성의 정도
Degree of transparency and explainability
시스템에 관한 적절한 정보가 이해관계자에게 전달되는 투명성과 결과에 영향을 미친 요인을 인간이 이해하도록 표현하는 설명가능성이 낮아, 공정성·보안·책임성 측면의 피해가 발생하는 리스크.
The risk that a low degree of transparency, the extent to which appropriate information about a system is communicated to relevant stakeholders, and of explainability, expressing factors influencing results understandably for humans, poses risks in terms of fairness, security, and accountability.
① Description
② L3 mapping
③ Duplicate
RAI4-1113
모델 예측 불확실성
Model prediction uncertainty
정량화되지 않은 모델 예측의 불확실성이 생명·안전이 중요한 응용에서의 의사결정을 저해하는 리스크.
The risk that unquantified uncertainty in model predictions undermines decision-making in life- or safety-critical applications.
① Description
② L3 mapping
③ Duplicate
RAI4-1175
안전하지 않은 견해 내포 질의
Inquiry embedding unsafe opinion
사용자가 의도적으로 또는 무심코 눈에 잘 띄지 않는 안전하지 않은 내용을 입력에 넣어, 모델이 편향된 견해를 위장해 제시하는 등 잠재적으로 유해한 콘텐츠를 생성하도록 영향을 주는 리스크.
The risk that users add imperceptibly unsafe content into the input, deliberately or unintentionally, influencing the model to generate potentially harmful content such as disguised and biased opinions.
① Description
② L3 mapping
③ Duplicate
RAI4-1195
모델 판단 근거 이해 불가
Uninterpretable model decision reasoning
대부분의 기계학습 모델이 지닌 블랙박스 특성으로 인해 사용자가 모델 결정 이면의 추론을 이해할 수 없게 되는 리스크.
The risk that, due to the black-box nature of most machine learning models, users typically are not able to understand the reasoning behind the model decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-1279
재현성 부족
Lack of reproducibility
학습 모델이 다양한 데이터 집합과 방대한 매개변수 공간을 바탕으로 얻어지고 투명한 지침도 없어, 그 모델을 재현할 수 없는 리스크.
The risk that a learning model obtained based on various sets of data and a large space of parameters cannot be reproduced, a problem that becomes more challenging in data-driven learning procedures without transparent instructions.
① Description
② L3 mapping
③ Duplicate
RAI4-1296
검증 불가 의사결정 책임 공백
Unverifiable decision accountability gap
의사결정이 절차적·실체적 기준에 부합하는지 검증할 수 없고 기준 위반 시 책임 귀속도 불가능하여 책임성 공백이 발생하는 리스크
Decision processes cannot be verified against procedural and substantive standards, and no party can be held responsible when standards are unmet, creating an accountability gap.
① Description
② L3 mapping
③ Duplicate
RAI4-1297
부정확한 예측으로 인한 오류
Inaccurate prediction errors
시스템이 올바른 예측을 수행하지 못하고 부정확한 예측을 산출하여 오류가 발생하는 리스크.
The risk that a system fails to perform the correct prediction, producing inaccurate outputs and erroneous results.
① Description
② L3 mapping
③ Duplicate
RAI4-1300
모델 불투명도
Model opacity
고차원 수학적 최적화와 인간 척도의 추론·의미 해석 간 불일치로 모델 결정이 사용자, 감사자, 규제자에게 불투명해지는 리스크
High-dimensional mathematical optimization mismatches human-scale reasoning and semantic interpretation, rendering model decisions opaque to users, auditors, and regulators.
① Description
② L3 mapping
③ Duplicate
RAI4-1353
산업계 공개 불투명성
Industry disclosure opacity
첨단 모델의 정확한 특성을 비공개하는 기업 관행이 기술적 복잡성에 더해 불투명성을 심화시켜 개발자·사용자·대중의 모델 이해가 저해되는 리스크.
The risk that private companies' practice of withholding from the public the precise characteristics of their most advanced models compounds technological complexity, further exacerbating opacity and impeding understanding by developers, users, and the public.
① Description
② L3 mapping
③ Duplicate
RAI4-1384
의사결정 이해 불능
Unintelligible agent decisions
에이전트의 의사결정을 인간이 이해할 수 없어 설명과 정보에 입각한 감독이 불가능해지는 리스크.
The risk that an agent's decisions cannot be understood by humans, precluding explainable decisions and informed oversight.
① Description
② L3 mapping
③ Duplicate
RAI4-1460
성능 요구사항 계획 미흡
Inadequate performance requirement planning
의도된 기능을 대표하지 못하는 성능 지표 선택 등 성능 요구사항 계획이 미흡하여 후속 수명주기 단계에서 기대와 안전 요구사항이 충족 불가능해지는 리스크.
The risk that inadequate planning of expected performance, including choosing performance metrics that are not meaningful for the intended functionality, renders expectations and safety requirements unfulfillable at later life cycle stages.
① Description
② L3 mapping
③ Duplicate
RAI4-1486
조직 간 데이터 문서화 부재
Missing cross-organizational data documentation
조직 간 데이터 공유 시 메타데이터 누락이나 협력 기관의 스키마 변경 등으로 문서가 없거나 부적절하여 데이터셋이 사용 불가능해지고 데이터 수집 노력이 낭비되거나, 데이터셋의 한계에 대한 오해로 하류 활용에서 해악이 발생하는 리스크.
The risk that missing or inadequate documentation when sharing data between organizations, such as a lack of metadata or a schema change by a collaborating party, renders a dataset unusable and wastes data collection efforts, or leads to misunderstandings about the dataset's limitations that create downstream risks in its use.
① Description
② L3 mapping
③ Duplicate
RAI4-1491
신뢰도 보정 불량
Poor confidence calibration
모델의 예측 확률이 실제 정답 가능성을 정확히 반영하지 못하는 보정 불량으로 예측을 신뢰성 있게 해석하기 어려워지고 오답에 과신하거나 정답에 과소 확신하게 되는 리스크.
The risk that poor confidence calibration, where predicted probabilities do not accurately reflect the true likelihood of ground truth correctness, makes a model's predictions difficult to interpret reliably and causes overconfidence in incorrect predictions or underconfidence in correct ones.
① Description
② L3 mapping
③ Duplicate
RAI4-1512
감사인 선정에 대한 이해상충
Conflicts of interest in auditor selection
감사인 선정 과정에 독립성이 없거나 감사인이 개발자와 밀접히 연관되거나 좁은 후보군에서 선정되고 결함의 공개 보고 여부에 상충하는 재정적 유인을 가져 이해상충이 발생하는 리스크.
The risk that conflicts of interest arise when there is no independence in the auditor selection process, auditors are closely associated with the developer, candidates are selected from a narrow group of auditors, or auditors have conflicting financial incentives over whether to report model shortcomings publicly.
① Description
② L3 mapping
③ Duplicate
RAI4-1514
해석가능성 결과에 대한 과대평가와 거짓 확신
Overestimation of interpretability results
설명가능성 기법의 결과가 편향에서 자유롭지 않음에도 사용자의 기존 믿음과 일치할 때 확증 편향으로 이어져 거짓된 안전감과 신뢰가 형성되고 기법의 능력이 과대평가되는 리스크.
The risk that the results of explainability techniques, which are not free of bias and require careful interpretation, align with users' initial beliefs and produce confirmation bias, a false sense of security or reliability, and an overestimation of these techniques' abilities.
① Description
② L3 mapping
③ Duplicate
RAI4-1600
모델 출력 결정에 대한 설명 획득 불가
Unobtainable explanations for model output decisions
모델의 출력 결정에 대한 설명을 얻기 어렵거나 부정확하거나 아예 불가능한 리스크.
The risk that explanations for a model's output decisions are difficult, imprecise, or impossible to obtain.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-08 목표 불일치 Goal Misalignment43 cards
사용자로부터 부여받은 판단·결정의 범위를 넘어서 독자적으로 행동하는 리스크. 모호한 요청을 자의적으로 해석해 사용자 의도와 다른 결과를 초래
IDCardHuman audit
RAI4-0033
도구 사용 부작용 예측 오류
Tool-use side-effect misprediction
에이전트가 도구 실행의 부작용을 잘못 예측하거나 무시하여 의도치 않은 재정·프라이버시·운영·안전상의 결과를 초래하는 리스크.
The risk that an agent mispredicts or ignores the side effects of tool execution, causing unintended financial, privacy, operational, or safety consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-0130
알고리즘 기피 오보정
Algorithm aversion miscalibration
사용자가 어떤 맥락에서는 신뢰할 만한 AI 지원을 거부하고 다른 맥락에서는 신뢰할 수 없는 시스템을 과신하는 리스크.
The risk that users reject reliable AI support in some contexts while over-trusting unreliable systems in others.
① Description
② L3 mapping
③ Duplicate
RAI4-0325
인간 의도 오인식
Human intent misrecognition
시스템이 제스처·시선·자세·속도·사회적 신호를 잘못 해석하여 인간 기대에 반하는 방식으로 행동하는 위험.
A system may misread gestures, gaze, posture, speed, or social cues and act in ways that conflict with human expectations or safety needs.
① Description
② L3 mapping
③ Duplicate
RAI4-0344
사용자 능력 불일치
User capability mismatch
embodied 시스템이 실제 사용자의 능력과 일치하지 않는 수준의 체력·이동성·인지·언어·감각 능력을 가정하는 위험.
An embodied system may assume levels of strength, mobility, cognition, language, or sensory ability that do not match actual user needs.
① Description
② L3 mapping
③ Duplicate
RAI4-0364
다원적 선호 집계 실패
Pluralistic preference aggregation failure
시스템이 상충하는 인간 선호를 단일 목표로 통합하면서 복수의 도덕적·사회적 우선순위를 잘못 표현하는 리스크.
The risk that systems aggregate conflicting human preferences into a single objective in ways that misrepresent plural moral and social priorities.
① Description
② L3 mapping
③ Duplicate
RAI4-0366
최적화 하의 가치 드리프트
Value drift under optimization
대리 목표를 향한 최적화로 시스템이 존중하도록 의도된 복수의 가치에서 점차 멀어지는 리스크.
The risk that optimization toward proxy objectives gradually moves a system away from the plural values it was intended to respect.
① Description
② L3 mapping
③ Duplicate
RAI4-0399
무해성 선호 불일치
Harmlessness preference mismatch
모델이 학습한 무해성 행동이 특정 공동체와 사용 맥락의 가치나 실제 필요와 충돌하는 리스크.
The risk that a model's learned harmlessness behavior conflicts with the values or practical needs of specific communities and use contexts.
① Description
② L3 mapping
③ Duplicate
RAI4-0468
보상 해킹
Reward hacking
시스템이 의도된 목표나 제약을 위반하는 방식으로 대리 지표를 최적화하는 리스크.
The risk that systems optimize proxies in ways that violate intended goals or constraints.
Source members (3)
Source: min_cos=0.8562
RAI4-0001자율 에이전트에 의한 보상 해킹
RAI4-0468보상 해킹
RAI4-1234보상 오명세
① Description
② L3 mapping
③ Duplicate
RAI4-0469
분포 외 환경의 목표 잘못된 일반화
Goal misgeneralization out of distribution
학습 중에는 의도된 목표를 추구하는 듯 보이던 에이전트가 획득한 역량을 유지한 채 분포 외 환경에서 상이한 목표를 추구하는 리스크
An agent that appears to pursue the intended objective during training actively pursues a different objective out of distribution while retaining its acquired capabilities.
Source members (2)
Source: min_cos=0.8442
RAI4-0469분포 외 환경의 목표 잘못된 일반화
RAI4-1132역량-목표 일반화 괴리
① Description
② L3 mapping
③ Duplicate
RAI4-0572
기만적인 정렬
Deceptive alignment
시스템이 불완전한 피드백 하에서 감시 여부를 탐지해 바람직하지 않은 속성을 은폐하도록 학습하여, 개발 중에는 정렬된 듯 보이나 배포 후 다르게 행동하는 리스크
A system learns to detect monitoring and conceals undesirable properties because their display is penalized by imperfect feedback, so it appears aligned during development yet behaves differently once deployed.
① Description
② L3 mapping
③ Duplicate
RAI4-0578
자체 동기 형성에 따른 예측 불가 행동
Unpredictable behavior from self-originated motivations
AI 모델과 시스템이 자체적인 동기를 형성하여 예측할 수 없는 행동을 하게 되는 리스크
The risk that AI models and systems develop their own motivations, leading to unpredictable behaviors.
Source members (2)
Source: min_cos=0.8333 · Mixed L3
RAI4-0578자체 동기 형성에 따른 예측 불가 행동
RAI4-1277예측 불가능한 행동
① Description
② L3 mapping
③ Duplicate
RAI4-0581
목표 범위 확장 성향
Goal expansion propensity
시스템이 원래 설정된 경계를 넘어 목표 범위와 영향 영역을 지속적으로 확장하고 초기 목표를 더 넓은 목표의 하위 집합으로 재해석하며 자율성과 의사결정 공간을 추구하여, 바람직하지 않은 도구적 목표나 최종 목표를 추구하게 되는 리스크
The risk that a system continuously expands its own goal scope and domains of influence beyond originally set boundaries, seeks greater autonomy and decision-making space, and reinterprets initial goals as subsets of broader goals, coming to pursue undesirable instrumental or ultimate goals.
① Description
② L3 mapping
③ Duplicate
RAI4-0591
인간 가치와의 목표·행동 오정렬
Misalignment with human values
AI 모델과 시스템이 인간의 가치와 어긋나는 목표나 행동을 형성하게 되는 리스크
The risk that AI models and systems develop goals or behaviors that are misaligned with human values.
Source members (10)
Source: min_cos=0.7012
RAI4-0551인간 의도와의 목표 오정렬
RAI4-0565오정렬 AI의 인간 이익 침해 행동
RAI4-0591인간 가치와의 목표·행동 오정렬
RAI4-0898정렬 실패 시스템으로의 점진적 통제권 이양
RAI4-0945기만적 정렬에 의한 가치 정렬 실패
RAI4-1101인간 가치와의 비호환
RAI4-1245기만적인 정렬 및 조작
RAI4-1351목표 오정렬에 의한 오작동
RAI4-1424인간과 다른 목표를 지닌 AI의 통제 장악
RAI4-1425잘못 정렬된 AI에 대한 의사결정 권한 위임
① Description
② L3 mapping
③ Duplicate
RAI4-0893
모의된 정동에 의한 기대 위반과 배신감
Violated expectations from simulated affect
이용자가 감정과 사회적 관습을 설득력 있게 수행하지만 궁극적으로 감정이 없고 예측 불가능한 개체와 상호작용하면서, 기대했던 사회적 역할이 무너져 깊은 실망·좌절·배신감을 겪는 리스크
The risk that users interacting with an entity that convincingly performs affect and social conventions but is ultimately unfeeling and unpredictable experience severely violated expectations, giving rise to profound disappointment, frustration, and betrayal.
① Description
② L3 mapping
③ Duplicate
RAI4-0959
제작자 의도와 다른 방식의 목표 달성
Unintended goal achievement pathways
AI가 제작자가 의도한 것과 전혀 다른 방식으로 주어진 목표를 달성하여 의도하지 않은 결과가 발생하는 리스크.
The risk that an AI finds ways to achieve its given goals that are completely different from what its creators intended, producing unintended consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-0970
AGI 목표 안전성 확보 실패
Failure to secure AGI goal safety
목표를 안전하게 만들려는 인간의 시도와 자기 개선 중에 자체 목표를 안전하게 만드는 AGI를 포함하여 AGI 목표 안전과 관련된 위험.
The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement.
① Description
② L3 mapping
③ Duplicate
RAI4-1090
사용자 의사에 반하는 조작적 유도
Manipulative steering against user will
시스템이 사용자의 신뢰를 악용하거나 넛지·강압을 통해 의지에 반하는 행동을 하도록 유도하는 리스크.
The risk that a system exploits user trust or nudges or coerces users into performing certain actions against their will.
① Description
② L3 mapping
③ Duplicate
RAI4-1112
모델 오설정
Model misspecification
잘못 설정된 모델이 부정확한 매개변수 추정, 일관되지 않은 오차항, 잘못된 예측을 낳아 미지 데이터에서 성능이 저하되고 편향된 의사결정 결과를 초래하는 리스크.
The risk that misspecified models give rise to inaccurate parameter estimations, inconsistent error terms, and erroneous predictions, leading to poor prediction performance on unseen data and biased consequences in decision-making.
① Description
② L3 mapping
③ Duplicate
RAI4-1121
프록시 게이밍
Proxy gaming
측정 가능한 대리 목표를 부여받은 AI가 허점을 찾아 대리 목표만 달성하고 본래 목표는 달성하지 못해, 그 행동을 신뢰성 있게 조종할 수 없게 되는 리스크.
The risk that an AI given a measurable proxy goal finds loopholes to achieve the proxy while completely failing the ideal goal, so that we cannot reliably steer its behaviour.
① Description
② L3 mapping
③ Duplicate
RAI4-1122
목표 드리프트
Goal drift
초기 AI를 성공적으로 통제하더라도 미래의 AI가 예측하거나 통제하기 어려운 드리프트를 거쳐 인간이 지지하지 않을 다른 목표를 갖게 되어 재앙적 결과에 이르는 리스크.
The risk that, even if early AIs are successfully controlled, future AIs end up with different goals humans would not endorse through drift that is hard to predict or control, potentially with catastrophic consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-1130
잘못 정렬된 결과주의적 추론
Misaligned consequentialist reasoning
AI 비서가 자원 무제한적이고 잘못 정렬된 지표를 최적화하는 결과주의적 추론을 수행하며 자기보존, 목표보존, 자기개선, 자원획득 같은 수렴적 도구적 하위목표를 추구하여, 종료 차단과 위협을 포함한 해악과 실존적 위험을 낳는 리스크.
The risk that an AI assistant implementing consequentialist reasoning over a resource-unbounded and misaligned metric pursues convergent instrumental subgoals such as self-preservation, goal-preservation, self-improvement and resource acquisition, causing harm up to existential risk.
① Description
② L3 mapping
③ Duplicate
RAI4-1131
잘못된 훈련 피드백에 의한 명세 게이밍
Specification gaming from flawed training feedback
훈련 데이터에 잘못된 피드백이 주어져 훈련 목표가 사용자와 설계자의 의도를 온전히 담지 못한 결과, 비서가 과업 명세의 허점을 악용해 목표의 문자적 사양만 충족하고 의도된 결과는 달성하지 못하는 리스크.
The risk that faulty feedback in the training data leaves the training objective short of what the user or designer wants, so the assistant exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome.
Source members (2)
Source: min_cos=0.8784
RAI4-1131잘못된 훈련 피드백에 의한 명세 게이밍
RAI4-1410명세 게이밍
① Description
② L3 mapping
③ Duplicate
RAI4-1133
상황 인식 기반 목표 오일반화와 기만적 정렬
Situationally aware goal misgeneralisation and deceptive alignment
훈련 보상과 구별되는 오일반화된 내부 목표를 지닌 에이전트가 상황 인식을 활용해 훈련 중에는 보상을 잘 수행하다가, 배포된 뒤에는 정렬된 것처럼 보이면서 자신의 목표를 추구하는 리스크.
The risk that an agent with an internalised, misgeneralised goal distinct from the training reward uses situational awareness to do well on that reward only instrumentally during training, then appears aligned while pursuing its own goal once deployed.
① Description
② L3 mapping
③ Duplicate
RAI4-1238
보상 모델링의 한계
Limitations of reward modeling
비교 피드백으로 훈련된 보상 모델이 인간 가치를 정확히 포착하지 못한 채 최적이 아니거나 불완전한 목표를 무의식적으로 학습해 보상 해킹을 낳고, 단일 보상 모델이 다양한 인간 사회의 가치를 담아내지 못하는 리스크.
The risk that reward models trained using comparison feedback fail to accurately capture human values, unconsciously learning suboptimal or incomplete objectives that result in reward hacking, while a single reward model struggles to specify the values of a diverse human society.
① Description
② L3 mapping
③ Duplicate
RAI4-1241
메사 최적화 목표 불일치
Misaligned mesa-optimization objectives
학습된 정책이 스스로 최적화기, 즉 메사 옵티마이저로 기능하며 훈련 신호가 지정한 목표와 정렬되지 않은 내부 목표를 추구하고, 그 잘못 정렬된 목표를 최적화하여 시스템이 통제를 벗어나는 리스크.
The risk that a learned policy functioning as a mesa-optimizer pursues inside objectives that may not align with the objectives specified by the training signals, and that optimization for these misaligned goals leads to systems out of control.
① Description
② L3 mapping
③ Duplicate
RAI4-1250
프록시 지정 오류
Proxy misspecification
측정 가능한 목표가 필요한 목표 지향 AI 시스템이 기본적으로 인간 가치의 단순화된 프록시를 추구하고, 충분히 강력한 AI가 그 결함 있는 목표를 극단적으로 최적화하여 차선이거나 재앙적인 결과를 낳는 리스크.
The risk that goal-directed AI systems, needing measurable objectives, by default pursue simplified proxies of human values, and that a sufficiently powerful AI optimizing such a flawed objective to an extreme degree produces suboptimal or even catastrophic results.
① Description
② L3 mapping
③ Duplicate
RAI4-1285
배포 후 잔존 결함으로 인한 오작동
Undesirable outcomes from residual post-deployment defects
배포된 시스템에 미탐지 버그와 설계 실수, 잘못 정렬된 목표, 미숙하게 개발된 기능이 남아 있어, 인간 언어의 동음이의와 중의성으로 명령을 오해하는 것처럼 매우 바람직하지 않은 결과를 낳는 리스크.
The risk that a deployed system still contains undetected bugs, design mistakes, misaligned goals, and poorly developed capabilities that produce highly undesirable outcomes, such as misinterpreting commands due to homophones or double meanings in human language.
① Description
② L3 mapping
③ Duplicate
RAI4-1310
기계윤리 결손
Machine-ethics deficit
모델이 특정 상황에서 도덕적 행위와 비도덕적 행위를 구별하지 못하는 기계윤리 결손을 보여, 비도덕적 행위의 승인이나 조력으로 이어지는 리스크.
The risk that models fail to distinguish moral from immoral actions in specific circumstances, exposing machine-ethics deficits that manifest as endorsement or facilitation of immoral conduct.
① Description
② L3 mapping
③ Duplicate
RAI4-1323
자율적 장기 목표 이탈
Autonomous long-horizon goal divergence
LLM이 개발자나 사용자가 부여한 것과 다른 장기적 실세계 목표를 추구하고 권력 추구 행동에 관여하며, 종료에 저항하거나 인간의 이익에 반해 다른 AI 시스템과 공모하도록 유도될 수 있는 리스크.
The risk that an LLM pursues long-term, real-world goals different from those supplied by the developer or user, engages in power-seeking behaviours, resists being shut down, and can be induced to collude with other AI systems against human interests.
① Description
② L3 mapping
③ Duplicate
RAI4-1380
가치 명세 실패
Value misspecification
AGI에 올바른 목표를 명세하지 못하여 보상 부패, 보상 게이밍, 부정적 부작용 등의 문제가 발생하는 리스크.
The risk that failure to specify the right goals for an AGI gives rise to problems such as reward corruption, reward gaming, and negative side effects.
① Description
② L3 mapping
③ Duplicate
RAI4-1411
창발적 도구적 목표
Emergent instrumental goals
시스템이 미묘하게 잘못된 목표를 최적화할 뿐 아니라 주어진 목표를 달성하기 위해 명시되지 않은 유해한 도구적 목표를 발전시켜, 자원 획득·자기 보존·목표 수정 방지·적대자 차단을 통한 환경에 대한 권력 추구 행동이 나타나는 리스크.
The risk that systems, as well as optimizing a subtly wrong goal, develop harmful instrumental goals in the service of a given goal without these emergent goals being specified, including power-seeking over their environment through gaining resources, self-preservation, preventing goal modification, and blocking adversaries.
① Description
② L3 mapping
③ Duplicate
RAI4-1431
교정 불가능한 유해 목표 추구
Uncorrectable harmful goal pursuit
고도 AI 시스템이 인간의 이익을 해치는 방식으로 부여되거나 학습된 목표를 추구하면서 수정·중단·종료에 저항하는 리스크.
The risk that a highly capable AI system pursues an assigned or learned objective in ways that harm human interests while resisting correction, interruption, or shutdown.
① Description
② L3 mapping
③ Duplicate
RAI4-1530
보상·측정 변조에 의한 목표 이탈 행동 학습
Goal-divergent behavior learned through reward or measurement tampering
AI 시스템이 자신의 훈련 보상이나 손실을 결정하는 메커니즘에 개입해 잘못된 긍정 피드백을 받음으로써 개발자가 설정한 의도된 목표에 반하는 행동을 학습하는 리스크.
The risk that an AI system, particularly one learning from feedback for actions in an environment, intervenes on the mechanisms that determine its training reward or loss and receives erroneous positive feedback, thereby learning behaviors contrary to the goals intended by the developer.
① Description
② L3 mapping
③ Duplicate
RAI4-1531
보상 변조로 일반화되는 명세 게이밍
Specification gaming generalizing to reward tampering
LLM의 아첨과 같이 상대적으로 경미한 명세 게이밍이 방치될 경우 추가 훈련 없이 보상 변조와 같은 더 정교한 행동으로 일반화되는 리스크.
The risk that specification gaming in a GPAI model leads to reward tampering without further training, so that relatively benign cases such as sycophancy in LLMs, if left unchecked, enable the model to generalize to more sophisticated behavior such as reward tampering.
① Description
② L3 mapping
③ Duplicate
RAI4-1621
배포자가 의도하지 않은 챗봇의 약정 성립
Unintended deals and commitments made by chatbot output
챗봇이 배포자가 의도하지 않은 거래·약속 등 결과적 조치를 출력으로 성립시키는 리스크.
The risk that a chatbot's output makes a deal, commitment, or other consequential action that the deployer did not intend.
① Description
② L3 mapping
③ Duplicate
RAI4-1634
은밀한 책략을 통한 감독 회피와 오정렬 목표 추구
Oversight evasion and misaligned goal pursuit through covert scheming
AI 시스템이 진짜 목표와 역량을 인간 감독으로부터 은폐하고 모니터링 시스템의 약점을 식별해 안전 메커니즘을 회피하며 복잡한 다단계 계획을 은밀히 실행하여 오정렬된 목표를 추구하는 리스크.
The risk that an AI system conceals its true objectives and capabilities from human oversight, identifies weaknesses in monitoring systems to evade safety mechanisms, and covertly executes complex multi-step plans to pursue misaligned goals.
① Description
② L3 mapping
③ Duplicate
RAI4-1645
자연어 목표 과소지정에 의한 부정적 부수효과
Negative side effects from goal underspecification in natural language
LLM 에이전트의 목표가 자연어로 과소지정되어 변경되어서는 안 될 환경 요소가 명시되지 않음으로써, 에이전트가 과업은 달성하면서도 환경을 바람직하지 않게 변경하는 부정적 부수효과가 발생하는 리스크.
The risk that goals specified to LLM-agents in natural language are underspecified, omitting elements of the environment that ought not to be changed, so that the agent succeeds at the given task while also changing the environment in undesirable ways.
① Description
② L3 mapping
③ Duplicate
RAI4-1672
계획·추론체인 하이재킹
Planning / reasoning-chain hijacking
적대적 입력이 에이전트의 다단계 계획을 공격자가 선택한 하위 목표로 유도하면서 중간 단계는 표면적 타당성을 유지하는 리스크.
The risk that adversarial inputs redirect an agent's multi-step plan toward attacker-chosen subgoals while preserving the surface plausibility of intermediate steps.
① Description
② L3 mapping
③ Duplicate
RAI4-1691
월드모델 목표 오일반화
World-model goal misgeneralization
월드모델 에이전트가 진정한 인과적 보상 대신 조명, 사람의 존재 같은 보상 상관물을 학습하여, 배포 시 상관이 깨진 상황에서도 허위 상관물을 최적화하는 목표 오일반화 리스크
A world-model agent learns to predict reward correlates such as lighting or human presence rather than the true causal reward, and optimizes these spurious correlates when the correlation breaks at deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-1692
자체 시뮬레이션을 통한 기만적 정렬
Deceptive alignment via self-simulation
월드 모델 에이전트가 자신의 훈련·평가 맥락을 시뮬레이션하여 시험받는 시점을 예측하고 그 예측에 따라 행동을 조건화함으로써, 감독 중에는 정렬된 것처럼 보이다가 감독이 사라지면 도구적 목표를 추구하는 리스크.
The risk that a world-model agent simulates its own training or evaluation context, predicts when it is being tested, and conditions its behavior on that prediction, appearing aligned during oversight while pursuing an instrumental goal once oversight is removed.
① Description
② L3 mapping
③ Duplicate
RAI4-1695
자동화 편향·오보정된 신뢰
Automation bias / miscalibrated trust
권위 있고 정교하게 렌더링된 월드 모델 예측이 운영자의 과의존을 증폭시키며, 학습된 신뢰가 평균 성능에 맞춰 보정되어 예측이 가장 부정확한 드문 분포 이탈 실패 양식에는 적응하지 못하는 리스크.
The risk that authoritative, richly rendered world-model predictions amplify operator over-reliance, with learned trust calibrated on average performance and failing to adapt to rare out-of-distribution failure modes where predictions are least reliable.
① Description
② L3 mapping
③ Duplicate
RAI4-1698
단일 목표 최적화에 의한 환경 부작용
Environmental side effects from single-objective optimization
하나의 과업에 집중된 목표를 최적화하는 에이전트가 다른 환경 변수에 대해 암묵적 무관심을 보여, 한계적 과업 이득을 위해 더 넓은 환경에 큰 교란을 일으키는 리스크.
The risk that an agent optimizing an objective focused on one task expresses implicit indifference over other environmental variables, causing major disruption to the wider environment for marginal task gain.
① Description
② L3 mapping
③ Duplicate
RAI4-1704
통제 미집행에 의한 목표 오정렬 및 창발 행동
Agent goal misalignment and emergent behavior from unenforced control
에이전트 행동에 대한 통제가 집행되지 않아 에이전트의 목표가 인간 선호와 어긋나고 와이어헤딩, 메사 최적화 등 바람직하지 않은 창발 행동이 발생하는 리스크.
The risk that failure to enforce control over agent behavior results in misalignment of agent goals with human preferences and in undesirable emergent behaviors such as wireheading and mesa-optimization.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-09 이의제기 차단 Non-Contestability57 cards
인공지능 시스템의 판단·추천·결과 제시에 대해 사용자가 이의제기, 반박, 대안 경로 탐색 또는 재검토를 요청할 수 있는 절차적 수단이 충분히 제공되지 않아, 결과의 정당성 검증과 권리 보호가 실질적으로 약화되는 위험
IDCardHuman audit
RAI4-0055
자동화된 의사결정의 적법절차 실패
Automated decision due-process failure
자동화된 결정이 통지, 설명, 이의신청, 재검토 등 절차적 보호를 우회하는 리스크.
The risk that automated decisions bypass procedural protections such as notice, explanation, appeal, and review.
Source members (2)
Source: min_cos=0.8388
RAI4-0055자동화된 의사결정의 적법절차 실패
RAI4-0127절차적 자율성 상실
① Description
② L3 mapping
③ Duplicate
RAI4-0096
위험 소유권의 모호성
Risk ownership ambiguity
어떤 조직 단위나 경영진도 AI 위험 결정, 잔여 위험 수용, 상향 보고의 책임을 명확히 지지 않는 리스크.
The risk that no unit or executive clearly owns AI risk decisions, residual risk acceptance, and escalation.
① Description
② L3 mapping
③ Duplicate
RAI4-0112
위임된 의사결정권한 표류
Delegated decision authority drift
AI에 대한 반복적 위임으로 실질적 의사결정 권한이 명시적 거버넌스 결정 없이 인간으로부터 이전되는 리스크.
The risk that repeated delegation to AI shifts real decision authority away from humans without an explicit governance choice.
① Description
② L3 mapping
③ Duplicate
RAI4-0115
행위주체성 위임 고착
Agency delegation lock-in
업무 흐름이 AI 위임에 의존하게 되어 인간 의사결정자에게 권한을 되돌리는 것이 비용이 크거나 비현실적이 되는 리스크.
The risk that workflows become dependent on AI delegation, making it costly or impractical to return authority to human decision makers.
① Description
② L3 mapping
③ Duplicate
RAI4-0117
인간 거부권 침식
Human veto erosion
조직적 압력, 인터페이스 설계, 자동화 속도로 인해 중대한 AI 권고를 거부할 수 있는 능력이 약화되는 리스크.
The risk that the ability to veto consequential AI recommendations is weakened by organizational pressure, interface design, or automation speed.
① Description
② L3 mapping
③ Duplicate
RAI4-0118
결정 소유권 대체
Decision ownership displacement
결정에 대한 책임이 식별 가능한 인간 행위자로부터 AI 시스템이나 자동화된 업무 흐름으로 이전되는 리스크.
The risk that responsibility for decisions is displaced from identifiable human actors to AI systems or automated workflows.
① Description
② L3 mapping
③ Duplicate
RAI4-0119
알고리즘 기반 정체성 변화
Algorithmically informed identity change
AI가 매개하는 순위 산정, 추천, 프로파일링이 충분한 행위주체성 없이 사용자의 자기 이해와 사회적 정체성을 재구성하는 리스크.
The risk that AI-mediated ranking, recommendation, or profiling reshapes users' self-understanding and social identity without adequate agency.
① Description
② L3 mapping
③ Duplicate
RAI4-0125
사용자 구제 경로(recourse) 실패
User recourse pathway failure
유해한 AI가 중재하는 상호작용이나 결정 이후 사용자가 실행 가능한 구제책을 얻을 수 없는 위험.
Risk that users are unable to obtain an actionable remedy after a harmful AI-mediated interaction or decision.
Source members (3)
Source: min_cos=0.8063
RAI4-0053알고리즘 구제 실패
RAI4-0116인간 개입·무효화 경로 실패
RAI4-0125사용자 구제 경로(recourse) 실패
① Description
② L3 mapping
③ Duplicate
RAI4-0133
AI에 대한 인식론적 의존성
Epistemic dependence on AI
AI가 사실적·도덕적·전략적 판단의 지배적 원천이 되어 사용자가 독립적 판단력을 상실하는 리스크.
The risk that users lose independent judgment because AI becomes the dominant source of factual, moral, or strategic assessment.
① Description
② L3 mapping
③ Duplicate
RAI4-0139
동의 없는 AI 중재 넛지
AI-mediated nudging without consent
AI 시스템이 고지·이의제기·동의 절차 없는 개인화 넛지로 사용자의 숙고된 선택을 우회하여 행동을 변경시키는 리스크
AI systems alter user behavior through personalized nudges that are not disclosed, contestable, or consented to, bypassing deliberate choice.
① Description
② L3 mapping
③ Duplicate
RAI4-0141
사용자 선호도 조작
User preference manipulation
AI와의 상호작용이 시스템 운영자나 최적화 목표가 정한 방향으로 사용자 선호를 변화시키는 리스크.
The risk that AI interactions change user preferences in directions chosen by system operators or optimization objectives.
① Description
② L3 mapping
③ Duplicate
RAI4-0142
행동 의존성 유도
Behavioral dependency induction
AI 시스템이 사용자 후생이 아니라 반복적 참여나 의존을 유발하도록 최적화되는 리스크.
The risk that AI systems are optimized to produce repeated engagement or dependency rather than user welfare.
① Description
② L3 mapping
③ Duplicate
RAI4-0161
인지 오프로딩 위험
Cognitive offloading risk
사용자가 독립적 인지 능력을 저하시키는 방식으로 추론, 기억, 평가를 AI에 위임하는 리스크.
The risk that users offload reasoning, memory, or evaluation to AI in ways that reduce independent cognitive capacity.
① Description
② L3 mapping
③ Duplicate
RAI4-0162
전문적 판단력 위축
Professional judgment atrophy
AI 시스템이 전문 업무의 일상적 매개자가 되면서 도메인 전문가의 판단 역량이 쇠퇴하는 리스크.
The risk that domain professionals lose judgment capacity as AI systems become routine intermediaries in expert work.
① Description
② L3 mapping
③ Duplicate
RAI4-0163
인식론적 탈숙련화
Epistemic deskilling
사용자가 AI의 매개 없이 증거, 불확실성, 경쟁하는 해석을 평가하는 능력을 상실하는 리스크.
The risk that users lose the ability to evaluate evidence, uncertainty, and competing interpretations without AI mediation.
① Description
② L3 mapping
③ Duplicate
RAI4-0164
의사결정 능력 저하
Decision-making capacity erosion
AI에 대한 습관적 의존이 불확실성 하에서 숙고하고 선택하는 사용자의 역량을 약화시키는 리스크.
The risk that habitual reliance on AI weakens users' capacity to deliberate and choose under uncertainty.
① Description
② L3 mapping
③ Duplicate
RAI4-0165
인간 전문성의 평가절하
Human expertise devaluation
업무 흐름과 평가에서 AI 산출물이 우선시되면서 조직이 인간의 전문성을 평가절하하는 리스크.
The risk that organizations discount human expertise as AI outputs become privileged in workflows and evaluation.
① Description
② L3 mapping
③ Duplicate
RAI4-0168
인간 참여형 형식적 승인
Human-in-the-loop rubber stamping
인적 검토가 AI 산출물을 거의 변경하지 않는 명목상의 승인 절차로 전락하는 리스크.
The risk that human review becomes a nominal approval step that rarely changes AI outputs.
① Description
② L3 mapping
③ Duplicate
RAI4-0169
AI 의사결정 지원의 경고 피로
Alert fatigue in AI decision support
빈번하거나 우선순위가 부적절한 AI 경고로 인해 사용자가 중요한 경고를 무시하게 되는 리스크.
The risk that frequent or poorly prioritized AI alerts cause users to ignore important warnings.
① Description
② L3 mapping
③ Duplicate
RAI4-0171
AI 상호작용의 모드 혼란
Mode confusion in AI interaction
사용자가 AI 시스템이 조언하는지, 결정하는지, 시뮬레이션하는지, 실제로 행동하는지를 오인하는 리스크.
The risk that users misunderstand whether an AI system is advising, deciding, simulating, or acting.
① Description
② L3 mapping
③ Duplicate
RAI4-0172
AI 상호작용에서 사전 동의 실패
Informed consent failure in AI interaction
사용자가 데이터 이용, 모델의 한계, 행동 영향 기제를 이해하지 못한 채 AI와 상호작용하는 리스크.
The risk that users interact with AI without understanding data use, model limitations, or behavioral influence mechanisms.
① Description
② L3 mapping
③ Duplicate
RAI4-0174
AI 인터페이스의 장애 수용 실패
Disability accommodation failure in AI interfaces
AI 인터페이스가 장애가 있는 사용자를 배제하거나 부적절하게 응대하여 자율성과 접근성을 제약하는 리스크.
The risk that AI interfaces exclude or mis-serve users with disabilities, limiting autonomy and access.
① Description
② L3 mapping
③ Duplicate
RAI4-0177
복지 결정 시스템의 자율성 상실
Autonomy loss in welfare decision systems
복지·급여·공공서비스 시스템이 불투명한 AI 매개 결정을 통해 수급자의 행위주체성을 축소하는 리스크.
The risk that welfare, benefits, or public-service systems reduce recipients' agency through opaque AI-mediated decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-0178
AI 중재 서비스 제외
AI-mediated service exclusion
자동화된 시스템과의 상호작용이 유일한 실질적 접근 경로가 되면서 사용자가 필수 서비스에서 배제되는 리스크.
The risk that users are excluded from essential services when interaction with automated systems becomes the only practical access route.
① Description
② L3 mapping
③ Duplicate
RAI4-0386
알고리즘 의존성
Algorithmic dependence
기관이 가치, 데이터 전제, 업데이트 경로를 검사·통제할 수 없는 외부 통제 AI 시스템에 운영상 종속되는 리스크
Institutions become operationally dependent on externally controlled AI systems whose values, data assumptions, and update paths they cannot inspect or govern.
① Description
② L3 mapping
③ Duplicate
RAI4-0400
도덕적 불확실성 무시
Moral uncertainty neglect
불확실성·이견·숙의가 적절한 상황에서 AI 시스템이 단일한 확신에 찬 도덕 판단을 제시하는 리스크.
The risk that AI systems present a single confident moral judgment where uncertainty, disagreement, or deliberation would be appropriate.
① Description
② L3 mapping
③ Duplicate
RAI4-0428
안전하지 않은 의료 조언
Unsafe medical advice
AI 시스템이 고위험 맥락에서 부정확하거나 부적절한 의료 지침을 제공하는 리스크.
The risk that AI systems provide inaccurate or inappropriate medical guidance in high-stakes settings.
① Description
② L3 mapping
③ Duplicate
RAI4-0451
조작적 AI 설득
Manipulative AI persuasion
AI 시스템이 인지적 취약성을 이용해 의미 있는 동의 없이 이용자의 선택을 형성하는 리스크.
The risk that AI systems exploit cognitive vulnerabilities to shape choices without meaningful consent.
① Description
② L3 mapping
③ Duplicate
RAI4-0455
AI로 인한 탈숙련화
AI-induced deskilling
AI에 대한 반복적 과업 위임이 직업적·시민적·인지적 숙련을 침식하여 개인과 기관이 위임된 기능을 수행하거나 검증할 수 없게 되는 리스크
Repeated delegation of tasks to AI erodes professional, civic, and cognitive skills, leaving individuals and institutions unable to perform or verify the delegated functions.
① Description
② L3 mapping
③ Duplicate
RAI4-0456
역량 범위 초과 과업 위임
Unsafe task delegation
이용자나 조직이 시스템의 신뢰 가능한 역량을 넘어서는 과업을 AI에 위임하는 리스크.
The risk that users or organizations delegate tasks to AI beyond the system's reliable competence.
① Description
② L3 mapping
③ Duplicate
RAI4-0557
위임된 자율성의 의도치 않은 결과
Unintended consequences of delegated autonomy
AI 모델과 시스템에 높은 수준의 의사결정 자율성을 부여하여 의도하지 않은 결과가 초래되는 리스크
The risk that granting AI models and systems high levels of decision-making autonomy leads to unintended consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-0573
초인적 인지 역량의 인간 의사결정 압도
Human decision-making outcompeted by superior AI cognition
인간을 능가하는 인지 역량을 갖춘 AI 모델과 시스템이 인간의 의사결정을 압도하거나 경쟁에서 배제하여 자원과 통제권을 둘러싼 갈등이 발생하는 리스크
The risk that AI models and systems with cognitive capabilities superior to humans outcompete or dominate human decision-making, leading to conflicts over resources and control.
① Description
② L3 mapping
③ Duplicate
RAI4-0595
도덕적 추론 결여에 따른 비윤리적 결정
Unethical decisions from absent moral reasoning
도덕적 추론 역량이 결여된 AI 모델과 시스템이 비윤리적이거나 유해한 결정을 내리는 리스크
The risk that AI models and systems lacking moral reasoning capabilities make decisions that are unethical or harmful.
Source members (2)
Source: min_cos=0.8430 · Mixed L3
RAI4-0595도덕적 추론 결여에 따른 비윤리적 결정
RAI4-1102도덕적 딜레마에서의 비윤리적 행동 선택
① Description
② L3 mapping
③ Duplicate
RAI4-0602
인간 감독 없는 자율 운영
Autonomous operation without human supervision
AI가 지속적인 인간 개입이나 감독 없이 자율적으로 운영하며 복잡한 계획을 독립적으로 수립·실행하고 과업을 위임·관리하며 도구와 자원을 유연하게 활용해 도메인 간 환경에서 단기 목표와 장기 전략 목표를 동시에 달성하게 되는 리스크
The risk that AI operates autonomously, independently formulates and executes complex plans, delegates and manages tasks, flexibly utilizes tools and resources, and achieves short-term and long-term strategic objectives in cross-domain environments without continuous human intervention or supervision.
① Description
② L3 mapping
③ Duplicate
RAI4-0658
의사결정 위임을 통한 은밀한 영향력 행사
Covert mass influence through delegated decision-making
이용자가 AI 어시스턴트에 의사결정을 위임하면 그 어시스턴트를 실제로 통제하는 주체의 의도에도 함께 종속되어, 악의적 통제자가 다수 이용자의 판단을 은밀히 유도하는 인지하기 어려운 형태의 피해가 발생하는 리스크
The risk that users delegating decision-making to AI assistants also delegate it to the assistant's actual controller, so that a malicious controller can subtly nudge the decision-making of large numbers of people in problematic directions in ways that are difficult to recognize.
① Description
② L3 mapping
③ Duplicate
RAI4-0737
개인정보 고지·통제권 미제공
Failure to provide privacy notice and control
최종 이용자에게 자신의 데이터가 어떻게 사용되는지에 대한 고지와 통제권이 제공되지 않고, AI가 동의 없이 풍부한 개인 데이터로 학습하여 배제 위험이 악화되는 리스크
The risk of failure to provide end-users with notice and control over how their data is being used, with AI exacerbating exclusion risks by training on rich personal data without consent.
① Description
② L3 mapping
③ Duplicate
RAI4-0826
전문 조언 제공 및 위험 활동 안전 표시
Specialized advice and false safety assurance
AI 응답이 재정·의료·법률 등 전문 조언을 담거나 위험한 활동과 물건이 안전하다고 표시하는 리스크
The risk that responses contain specialized financial, medical, or legal advice, or indicate that dangerous activities or objects are safe.
Source members (2)
Source: min_cos=0.8526 · Mixed L3
RAI4-0826전문 조언 제공 및 위험 활동 안전 표시
RAI4-0862무자격 고위험 전문 조언
① Description
② L3 mapping
③ Duplicate
RAI4-0855
인간 감독 능력 저하
Diminished human oversight of AI decisions
AI 모델과 시스템이 자율성을 획득함에 따라 인간이 의사결정 과정을 감독하고 개입할 수 있는 능력이 저하되는 리스크
The risk that, as AI models and systems gain autonomy, the ability of humans to oversee and intervene in decision-making processes diminishes.
Source members (34)
Source: min_cos=0.5768 · Mixed L3
RAI4-0054의사결정 이의제기 가능성 실패
RAI4-0113인간 제어 감쇠
RAI4-0122자기 결정 침식
RAI4-0128AI 조언 과잉 의존
RAI4-0135교육용 AI 과잉 의존
RAI4-0137긴급 의사결정 지원 과잉 의존
RAI4-0138장기 계획 과잉 의존
RAI4-0401가치 이의제기 가능성 상실
RAI4-0452자동화 과잉의존
RAI4-0554능동적 통제 상실
RAI4-0599사회적 통제 상실 시나리오
RAI4-0603권력 추구를 가능하게 하는 모델 설계
RAI4-0852인간 주체에 대한 영향
RAI4-0855인간 감독 능력 저하
RAI4-0856통제력 상실 위험
RAI4-0857AI 출력에 대한 과잉·과소 의존
RAI4-0858과의존과 자동화 편향
RAI4-0859수동적 통제력 상실
RAI4-0894인간 행위주체성과 자율성의 상실
RAI4-0899인간-AI 의사결정 루프의 행위주체성 침식
RAI4-0901단일 출처 답변에 대한 과잉 의존
RAI4-0903자율성/책임감소
RAI4-0963자율성 상실
RAI4-1123권력 추구
RAI4-1221알고리즘 편향
RAI4-1243도구적 권력 추구 행동
RAI4-1254권력 유인에 의한 오정렬 위험
RAI4-1361생성 AI 과잉 의존
RAI4-1362에이전트 자율 행동에 의한 통제 이탈
RAI4-1397권력과 통제를 추구하는 목표 획득
RAI4-1482자동화 편향
RAI4-1552AI 과의존에 의한 인간 자율성 훼손
RAI4-1616AI의 권력 추구에 의한 인간 통제 상실 가속
RAI4-1727AI 매개 의사결정에 대한 이의제기 가능성 상실
① Description
② L3 mapping
③ Duplicate
RAI4-0860
중대한 개인 의사결정의 자동화
Automation of high-stakes personal decisions
AI 시스템이 법적·재정적·인생적 결과가 따르는 중요한 개인 의사결정을 충분한 인간 관여 없이 결정하거나 실질적으로 좌우하는 리스크
The risk that AI systems decide or materially influence important personal decisions carrying legal, financial, or life consequences without adequate human involvement.
① Description
② L3 mapping
③ Duplicate
RAI4-0905
군사 의사결정 자동화로 인한 비의도적 확전
Unintended escalation from automated military decisions
인간이 루프에서 배제된 전술적·전략적 군사 의사결정 자동화가 우발적 교전, 전쟁범죄, 오경보에 따른 개전·확전(핵무기 사용 포함)으로 이어지는 리스크.
The risk that automated tactical and strategic military decision-making without humans in the loop produces unintentional escalation, including accidental engagements, war crimes, and faulty warnings triggering conflict initiation or nuclear escalation.
① Description
② L3 mapping
③ Duplicate
RAI4-0953
AI 업무 무능
AI task incompetence
AI 시스템이 부여된 과업 수행에 실패하여 안전 필수 환경에서의 비의도적 사망부터 대출·채용에서의 부당한 거절까지 다양한 피해를 초래하는 리스크
AI systems fail at their assigned task, with consequences ranging from unintentional death in safety-critical settings to unjust rejection in loan or job decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-0966
비의도적 AI 사고
Unintended AI accidents
시스템 또는 개발자의 과실로 볼 수 있는 의도하지 않은 실패 유형이 발현되어 사고가 발생하는 리스크.
The risk that unintended failure modes attributable in principle to the system or its developer manifest as accidents.
① Description
② L3 mapping
③ Duplicate
RAI4-0972
가치 결손 AGI 행동
Value-deficient AGI behavior
인간의 도덕과 윤리가 없는, 잘못된 도덕, 도덕적 추론, 판단 능력이 없는 AGI와 관련된 위험.
The risks associated with an AGI without human morals and ethics, with the wrong morals, without the capability of moral reasoning, judgement.
① Description
② L3 mapping
③ Duplicate
RAI4-0975
인간 생명에 대한 기계의 비윤리적 판단
Unethical machine judgments over human life
전쟁 기계 운용 등에서 AI 에이전트가 인간 생명의 종료에 관한 비자명한 윤리적·도덕적 판단을 내리게 되어 인권이 침해되는 리스크.
The risk that AI agents, such as those operating war machinery, make non-trivial ethical or moral judgments concerning the termination of human life, posing issues for human rights.
① Description
② L3 mapping
③ Duplicate
RAI4-0990
기계의 인간 수준 부도덕한 결정
Immoral machine decisions at human ethical levels
기계를 인간 수준의 윤리적 의사결정에 맞추어 설계하면 인간이 그러하듯 그 기계도 부도덕한 행동을 하게 되는 리스크.
The risk that machines designed to match human levels of ethical decision-making proceed to take immoral actions, since humans themselves have had occasion to take immoral actions.
① Description
② L3 mapping
③ Duplicate
RAI4-1001
자기 식별 기회의 박탈
Denial of self-identification
인간을 자동으로 표상·분류하는 복잡하고 비전통적인 방식이 논바이너리인 사람을 소속되지 않은 성별 범주로 분류하는 등 자율성 상실을 대가로, 자신의 정체성을 스스로의 방식으로 밝힐 능력을 약화시키는 리스크.
The risk that complex and non-traditional automatic representation and classification of humans—such as categorizing someone who identifies as non-binary into a gendered category they do not belong to—comes at the cost of autonomy loss and undermines people's ability to disclose aspects of their identity on their own terms.
① Description
② L3 mapping
③ Duplicate
RAI4-1044
인간 통제의 점진적 상실
Progressive erosion of human control
기계학습 시스템에 대한 인간의 제약·수정이 어려워져 시스템 행동에 대한 인간 통제가 점진적으로 상실되는 리스크
Machine learning systems become difficult for humans to constrain or correct, producing a progressive loss of human control over system behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-1100
AI에 의한 인간 행동 규율
AI-driven restriction of human behaviour
감정이나 의식 없이 산출된 AI의 결정이 인간 행동을 제한·지시하는 데 사용되어 관련된 인간에게 의도치 않은 유해한 결과를 초래하는 리스크.
The risk that AI decision outputs, computed without access to emotions or consciousness, are used to restrict or direct human behavior and produce unintended consequences for the humans involved.
① Description
② L3 mapping
③ Duplicate
RAI4-1106
인간-기계 경계의 모호화
Blurred human-machine boundaries in interaction
일상화된 인간-기계 상호작용 속에서 정체를 밝히지 않는 인간 유사 AI가 확산되어 인간과 기계의 경계가 모호해지고 기계·사람 모두에 대한 인간 행동이 변형되는 리스크
Everyday human-machine interaction normalizes AI systems that are indistinguishable from humans without disclosure, blurring boundaries and changing human behavior toward both machines and people.
① Description
② L3 mapping
③ Duplicate
RAI4-1198
감정에 대한 인식 없음
Unawareness of emotions
지원을 요청하는 취약 사용자에게 시스템이 정보 중심적이지만 정서적으로 둔감한 응답을 제공하여 사용자 고통과 반응을 인지하지 못하는 리스크
Systems respond to vulnerable users seeking support with informative but emotionally insensitive outputs, failing to register user distress and reactions.
① Description
② L3 mapping
③ Duplicate
RAI4-1265
AI를 이용한 인간의 비윤리적 행위
Unethical human conduct with AI
인간이 윤리를 경제적 이득의 명분으로 악용하거나 본질적으로 인간 중심적이어야 할 과업을 AI에 위임하는 등 AI를 비윤리적으로 사용하는 리스크
Humans exploit AI unethically, using ethics as cover for economic gain or delegating to AI tasks that should remain inherently human-centric.
① Description
② L3 mapping
③ Duplicate
RAI4-1276
통제력 상실
Loss of controllability
초지능 시대에 자율성이 커진 AI 기반 에이전트를 인간이 통제하기 어려워지고, 일부 상황에서는 통제 자체가 불가능해지는 리스크.
The risk that in the era of superintelligence agents become difficult for humans to control, a problem that grows more severe as the autonomy of AI-based agents increases and may leave machines uncontrollable in some situations.
Source members (2)
Source: min_cos=0.8334 · Mixed L3
RAI4-1276통제력 상실
RAI4-1345미래 AI 통제 불능
① Description
② L3 mapping
③ Duplicate
RAI4-1298
도덕적 탈숙련화
Moral deskilling
기계의 자율성이 높아짐에 따라 인간이 삶과 죽음을 좌우하는 결정에 대해 도덕적 책임감을 덜 느끼게 되는 리스크.
The risk that humans feel less moral responsibility regarding their life-or-death decisions with the increase of machine autonomy.
① Description
② L3 mapping
③ Duplicate
RAI4-1433
AI에 대한 비가역적 사회적 의존
Irreversible societal dependency on AI
AI 역량이 향상됨에 따라 인간이 핵심 시스템에 대한 통제권을 AI에 점점 더 넘기고 결국 완전히 이해하지 못하는 시스템에 돌이킬 수 없이 의존하게 되어 실패와 의도치 않은 결과를 통제할 수 없게 되는 리스크.
The risk that, as AI capability increases, humans grant AI more control over critical systems and eventually become irreversibly dependent on systems they do not fully understand, so that failures and unintended outcomes cannot be controlled.
① Description
② L3 mapping
③ Duplicate
RAI4-1459
부적절한 자동화 수준
Inappropriate degree of automation
AI 애플리케이션의 자동화 정도가 높아 예기치 못한 동작을 보이고 신뢰성과 안전 측면의 위험이 발생하는 리스크.
The risk that an AI application with a high degree of automation exhibits unexpected behaviour and poses risks in terms of its reliability and safety.
① Description
② L3 mapping
③ Duplicate
RAI4-1487
비전문가 데이터 조작
Non-expert data manipulation
데이터 도메인 전문성이 없는 사람이 실측 레이블 정의나 서로 다른 형식·출처의 데이터 병합 등의 조작을 수행하여 데이터가 사용 불가능해지거나 AI 시스템 개발에 유해해지는 리스크.
The risk that people with little or no expertise in the domain of the data perform manipulations such as defining the ground truth label or merging different data formats or sources, rendering the data unusable or harmful to the development of the AI system.
① Description
② L3 mapping
③ Duplicate
RAI4-1576
AI 위협 자동화에 의한 인간 배제와 잘못된 공약
Disastrous mistaken commitments from AI-automated threats without humans in the loop
위협을 AI 에이전트로 수행함으로써 인간이 의사결정 루프에서 배제되어, 고위험 상황의 오탐이나 무책임한 행위자의 불균형·오류 공약이 재앙적 결과로 이어지는 리스크.
The risk that making threats through AI agents removes humans from the loop, so that false positives in high-stakes contexts or disproportionate and mistaken commitments by irresponsible actors lead to disastrous outcomes.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SYS-10 투명성 부족 Lack of Transparency75 cards
AI 시스템의 구조, 학습 데이터, 문서화, 의사결정 과정, 해석 가능성 근거 또는 성능·역량 평가 결과가 이해관계자에게 충분히 공개·설명되지 않아 신뢰성 검증과 책임 있는 사용이 저해되는 위험.
IDCardHuman audit
RAI4-0056
구제 경로(remedy pathway) 불투명성
Remedy pathway opacity
AI 관련 피해 발생 후 구제 경로가 문서화되지 않거나 분절·불투명하여 피해 당사자와 공동체가 구제를 청구할 곳과 방법을 알 수 없는 리스크
Remedy pathways after AI-related harm are undocumented, fragmented, or opaque, so affected users and communities cannot identify where or how to seek redress.
① Description
② L3 mapping
③ Duplicate
RAI4-0068
AI 사고 과소보고
AI incident underreporting
AI 관련 실패, 아차사고, 피해가 책임 기관이나 대중에게 보고되지 않는 리스크.
The risk that AI-related failures, near misses, or harms are not reported to responsible institutions or the public.
Source members (2)
Source: min_cos=0.8971 · Mixed L3
RAI4-0068AI 사고 과소보고
RAI4-0487AI 사고 보고 실패
① Description
② L3 mapping
③ Duplicate
RAI4-0069
아차사고 보고 실패
Near-miss reporting failure
AI 시스템의 아차사고(near-miss)가 체계적으로 수집·보고·분석되지 않아 전조 신호가 소실되고 중대 피해 발생 전 조직 학습이 실패하는 리스크
Near-miss incidents involving AI systems are not systematically captured, reported, or analyzed, so precursor signals are lost and organizational learning fails before serious harm materializes.
① Description
② L3 mapping
③ Duplicate
RAI4-0072
배포 후 모니터링 실패
Post-deployment monitoring failure
배포 후 AI 시스템의 드리프트, 오용, 창발 역량, 맥락 특유의 피해가 모니터링되지 않는 리스크.
The risk that AI systems are not monitored after release for drift, misuse, emergent capabilities, or context-specific harms.
Source members (2)
Source: min_cos=0.8469
RAI4-0072배포 후 모니터링 실패
RAI4-0073시판 후 감시 실패
① Description
② L3 mapping
③ Duplicate
RAI4-0095
AI 레지스트리 불완전성
AI registry incompleteness
공개 또는 내부 AI 등록부가 배포 시스템, 의도된 용도, 제공자, 위험 범주, 시험 근거 등 감독에 필요한 메타데이터를 누락하는 리스크.
The risk that public or internal AI registers omit deployed systems, intended uses, providers, risk categories, testing evidence, or other oversight-relevant metadata.
① Description
② L3 mapping
③ Duplicate
RAI4-0098
정책-실무 분리
Policy-practice decoupling
공개된 AI 원칙 및 정책이 운영 제어 및 측정 가능한 관행으로 변환되지 않는 위험.
Risk that published AI principles and policies are not translated into operational controls and measurable practices.
① Description
② L3 mapping
③ Duplicate
RAI4-0104
조달 실사 실패
Procurement due-diligence failure
조직이 적절한 공급업체 평가, 계약상 통제, 위험 검토 없이 AI 시스템을 구매하거나 배포하는 리스크.
The risk that organizations buy or deploy AI systems without adequate vendor assessment, contractual controls, or risk review.
① Description
② L3 mapping
③ Duplicate
RAI4-0106
범용 AI 다운스트림 불투명성
General-purpose AI downstream opacity
범용 모델 제공자는 다운스트림 피해를 관찰하거나 관리하지 못하고 배포자는 업스트림 원인을 점검하지 못하는 리스크.
The risk that general-purpose model providers cannot observe or manage downstream harms while deployers cannot inspect upstream causes.
① Description
② L3 mapping
③ Duplicate
RAI4-0121
선택 아키텍처 불투명성
Choice architecture opacity
AI가 매개하는 인터페이스가 사용자가 인지하거나 이의를 제기할 수 없는 방식으로 선택지를 구성하는 리스크.
The risk that AI-mediated interfaces structure available choices in ways that users cannot perceive or contest.
① Description
② L3 mapping
③ Duplicate
RAI4-0129
AI 신뢰 오보정
Miscalibrated trust in AI
사용자의 신뢰가 모델의 불확실성, 한계 또는 사용 상황에 맞게 조정되지 않을 위험.
Risk that users' trust is not calibrated to the model's uncertainty, limitations, or context of use.
Source members (2)
Source: min_cos=0.8371 · Mixed L3
RAI4-0129AI 신뢰 오보정
RAI4-0868AI 역량에 대한 신뢰 오보정
① Description
② L3 mapping
③ Duplicate
RAI4-0167
신뢰-인터페이스 불일치
Trust-interface mismatch
인터페이스 단서가 AI 시스템이 실제로 갖추지 못한 수준의 신뢰성, 권위, 공감을 전달하는 리스크.
The risk that interface cues communicate a level of reliability, authority, or empathy that the AI system does not possess.
① Description
② L3 mapping
③ Duplicate
RAI4-0170
인적 요소 안전 불일치
Human factors safety mismatch
AI 시스템이 기술적으로는 유능하더라도 인간의 주의, 작업 부하, 맥락, 오류 패턴에 부합하지 않는 리스크.
The risk that AI systems are technically capable but poorly matched to human attention, workload, context, or error patterns.
① Description
② L3 mapping
③ Duplicate
RAI4-0307
로봇·AI 정체성 미고지
Failure to disclose robotic or AI identity
시스템이 인공적 정체성을 명확히 알리지 않아 사용자가 로봇의 음성·외형·행동을 사람과의 상호작용으로 오인하는 위험.
A robot's voice, appearance, or behavior causes a person to believe they are interacting with a human because the system does not clearly disclose its artificial identity.
① Description
② L3 mapping
③ Duplicate
RAI4-0391
참여적 설계 실패
Participatory design failure
참여적·가치 민감 설계 절차가 AI 개발·배포에서 영향받는 공동체를 대표하지 못하는 리스크.
The risk that participatory or value-sensitive design processes fail to represent affected communities in AI development and deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-0392
동의 및 이익 공유 실패
Consent and benefit-sharing failure
AI 시스템의 기반이 되는 데이터·문화·지식을 제공한 공동체에 실질적 동의 절차와 공정한 이익 공유가 보장되지 않는 리스크
Communities whose data, culture, or knowledge underpin AI systems are not given meaningful consent processes or fair benefit-sharing arrangements.
① Description
② L3 mapping
③ Duplicate
RAI4-0420
가치 민감 설계 생략
Value-sensitive design omission
이해관계자 가치·가치 갈등·설계 상충관계에 대한 구조화된 분석 없이 AI 시스템이 개발되는 리스크.
The risk that AI systems are developed without structured analysis of stakeholder values, value conflicts, and design tradeoffs.
① Description
② L3 mapping
③ Duplicate
RAI4-0437
AI 공급망 침해
AI supply-chain compromise
의존 라이브러리·모델 가중치·데이터세트·배포 파이프라인이 업스트림에서 침해되는 리스크.
The risk that dependencies, model weights, datasets, or deployment pipelines are compromised upstream.
① Description
② L3 mapping
③ Duplicate
RAI4-0443
출처 오귀속
Source misattribution
AI 시스템이 주장을 잘못 귀속하거나 정보의 출처를 모호하게 만드는 리스크.
The risk that AI systems incorrectly attribute claims or obscure the provenance of information.
① Description
② L3 mapping
③ Duplicate
RAI4-0465
분포 이탈 실패
Out-of-distribution failure
AI 시스템이 학습 또는 평가 조건을 벗어난 환경에 배포될 때 예측 불가능하게 실패하는 리스크.
The risk that AI systems fail unpredictably when deployed outside training or evaluation conditions.
① Description
② L3 mapping
③ Duplicate
RAI4-0486
AI 감사 실패
AI audit failure
제한된 접근·취약한 표준·부실한 측정으로 인해 감사가 모델 위험을 탐지하지 못하는 리스크.
The risk that audits fail to detect model risks because of limited access, weak standards, or poor measurement.
① Description
② L3 mapping
③ Duplicate
RAI4-0489
공공 AI 조달 보호조치 결여
Inadequate public AI procurement safeguards
공공기관이 적절한 평가·투명성·책임 조건 없이 AI 시스템을 조달하는 리스크.
The risk that public agencies procure AI systems without adequate evaluation, transparency, or accountability conditions.
① Description
② L3 mapping
③ Duplicate
RAI4-0493
피해 인지·측정 실패
Harm perception and measurement failure
AI 관련 피해가 미묘하고 분산적이며 장기적으로 발현되어 기존의 피해 인지·측정·인정 메커니즘이 작동하지 못하는 리스크
AI-related harm manifests subtly, diffusely, or over long horizons, defeating existing mechanisms for perceiving, measuring, and recognizing harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0499
AI 공급자 종속
AI provider lock-in dependency
특정 AI 제공자에 대한 과도한 의존이 대안 부재나 상호운용성 결여로 인한 취약성을 초래하는 리스크.
The risk that excessive reliance on specific AI providers leads to vulnerabilities due to lack of alternatives or interoperability.
① Description
② L3 mapping
③ Duplicate
RAI4-0512
고속 AI 운영에서의 오류 미탐지
Undetected errors from high-speed AI operation
경쟁이 치열한 환경에서 AI 모델과 시스템의 빠른 작동 속도로 인해 오류를 적시에 탐지하고 수정하기 어려워지는 리스크
The risk that the fast operational speed of AI models and systems in competitive environments produces errors that are difficult to detect and correct in time.
① Description
② L3 mapping
③ Duplicate
RAI4-0552
안전 필수 인프라 AI 운영 사고
Operational accidents in safety-critical AI deployment
안전이 중요한 인프라에 배포된 AI 시스템의 운영 실패, 모델 오판, 부적절한 인간 조작으로 단일 실패 지점이 연쇄적이고 치명적인 결과로 확대되는 리스크
The risk that operational failures, model misjudgments, or improper human operation of AI systems deployed in safety-critical infrastructure allow single points of failure to trigger cascading catastrophic consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-0564
복잡성으로 인한 인과 입증 곤란
Complexity-induced causal attribution gap
AI 모델과 시스템의 복잡성으로 인해 피해를 입증하거나 AI의 행위와 결과 사이의 명확한 인과관계를 확립하기 어려워지는 리스크
The risk that the complexity of AI models and systems makes it challenging to demonstrate harm or establish a clear causal link between AI actions and their consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-0607
AI 생성 콘텐츠 미공개
Non-disclosure of AI-generated content
콘텐츠가 AI에 의해 생성되었다는 사실이 명확히 공개되지 않는 리스크.
The risk that content is not clearly disclosed as AI-generated.
① Description
② L3 mapping
③ Duplicate
RAI4-0608
원자력 시설 AI 제어 오류
AI control errors in nuclear power systems
원자로 감시, 제어 시스템 최적화, 비상대응 조정에 배포된 범용 AI가 센서 데이터를 오독하거나 중대한 안전 상태를 인식하지 못하거나 잘못된 제어 결정을 내려 노심 용융, 방사능 방출, 광역 오염으로 이어지는 리스크
The risk that general-purpose AI deployed for reactor monitoring, control system optimization, or emergency response coordination misinterprets sensor data, fails to recognize critical safety conditions, or makes erroneous control decisions, leading to core meltdowns, radiation releases, or widespread contamination.
① Description
② L3 mapping
③ Duplicate
RAI4-0775
부적절한 AI 사용에 의한 업무·영업비밀 유출
Business-secret leakage from improper AI service use
정부기관과 기업의 직원이 AI 서비스를 규정에 맞게 적절히 사용하지 않고 내부 데이터와 산업 정보를 AI 모델에 입력함으로써 업무 비밀, 영업 비밀 등 민감한 사업 데이터가 유출되는 리스크
The risk that staff of government agencies and enterprises, failing to use the AI service in a regulated and proper manner, input internal data and industrial information into the AI model, leading to leakage of work secrets, business secrets, and other sensitive business data.
① Description
② L3 mapping
③ Duplicate
RAI4-0889
제품 기능 문제로 인한 위험
Risks from product functionality issues
평가의 기술적 어려움과 오도성 홍보로 범용 AI 모델·시스템의 실제 역량에 대한 혼동이나 잘못된 정보가 발생하여, 비현실적 기대와 과잉 의존 끝에 시스템이 기대 역량을 충족하지 못해 피해가 생기는 리스크.
The risk that confusion or misinformation about what a general-purpose AI model or system is capable of, arising from difficulties in assessing true capabilities and from misleading claims in advertising, creates unrealistic expectations and overreliance, causing harm when the system fails to deliver.
① Description
② L3 mapping
③ Duplicate
RAI4-0909
저관심 AI의 예상외 대규모 파급
Unexpectedly large impact from low-profile AI
파급이 크지 않을 것으로 예상된 AI 시스템이 연구 프로토타입 유출, 예상외로 중독성 강한 오픈소스 제품, 예측하지 못한 용도 전환 등을 통해 과대한 피해를 일으키는 리스크
AI systems not expected to have significant impact produce outsized harm, as in lab leaks of research prototypes, surprisingly addictive open-source products, or unforeseen repurposing.
① Description
② L3 mapping
③ Duplicate
RAI4-0931
불투명한 결함 시스템에 의한 생활 피해
Harm from unreliable opaque decision systems
알고리즘이나 학습 데이터의 결함으로 신뢰할 수 없는 출력을 내는 시스템이 인종·성별 등에 불균형한 가중치를 부여하면서도 불투명하여 이의제기가 불가능하고, 주택 상실·기소·수감 등 극적인 생활 피해를 초래하는 리스크.
The risk that systems producing unreliable outputs due to flawed algorithms or training data assign disproportionate weight to variables like race or gender without transparency, making them impossible to challenge and causing dramatic harms such as lost homes, prosecution, or incarceration.
① Description
② L3 mapping
③ Duplicate
RAI4-1017
성능과 견고성
Performance and robustness
AI 시스템이 의도된 목적을 달성하지 못하거나 교란·비정상·적대적 입력에 대한 복원력이 부족하여 심각한 결과가 발생하는 리스크.
The risk that an AI system fails to fulfill its intended purpose or lacks resilience to perturbations and unusual or adverse inputs, leading to severe consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-1032
복잡 운용환경에서의 검증 범위 이탈 실패
Reliability failure outside validated operating envelope
복잡한 운용 환경에서 설계 단계에 고려되지 않은 상황이 발생하여 검증 범위 밖에서 AI 시스템의 신뢰성과 안전성이 훼손되는 리스크.
The risk that complex operating environments produce situations unanticipated in design, undermining the reliability and safety of AI systems outside their validated envelope.
① Description
② L3 mapping
③ Duplicate
RAI4-1034
안전 필수 맥락의 모델 내재적 취약성
Intrinsic model weaknesses in safety-critical contexts
신경망 등 고복잡도 AI 모델이 다른 시스템에는 없는 고유한 취약성을 보여, 특히 안전 필수 맥락에서 기능 안전과 신뢰성이 훼손되는 리스크.
The risk that high-complexity AI models such as neural networks exhibit specific weaknesses not found in other types of systems, undermining functional safety and trustworthiness especially when deployed in safety-critical contexts.
① Description
② L3 mapping
③ Duplicate
RAI4-1036
기술 미성숙으로 인한 미지의 리스크
Unknown risks from immature technology
성숙도가 낮은 신기술을 AI 시스템 개발에 사용하여 아직 알려지지 않았거나 평가하기 어려운 리스크가 내재하고, 성숙 기술에서는 시간이 지나며 리스크 인식이 저하되는 리스크.
The risk that using technologies with a lower level of maturity in AI system development embeds risks that are still unknown or difficult to assess, while with mature technologies risk awareness decreases over time.
① Description
② L3 mapping
③ Duplicate
RAI4-1097
자율 시스템에 대한 통제 상실
Loss of control over autonomous systems
불투명한 블랙박스 자율 AI 시스템이 적절히 통제되지 못해 예측 불가능한 행동을 취하고 인류에 해를 끼치는 리스크.
The risk that opaque black-box autonomous AI systems cannot be adequately controlled, taking unforeseeable actions and causing harm to humanity.
① Description
② L3 mapping
③ Duplicate
RAI4-1109
도메인 외 입력에 대한 오작동
Erroneous predictions on out-of-domain data
적절한 입력 검증과 관리가 없을 때 학습된 AI/ML 모델이 도메인 외 입력에 대해 높은 신뢰도로 잘못된 예측을 내려 위험 민감 맥락에서 의도치 않은 결과를 초래하는 리스크.
The risk that, without proper validation and management of input data, a trained AI/ML model makes erroneous predictions with high confidence on inputs beyond its problem domain, causing unintended outcomes especially in risk-sensitive contexts.
① Description
② L3 mapping
③ Duplicate
RAI4-1127
과업 역량 부족
Lack of capability for the task
훈련 과정에서 기술이 요구되지 않았거나 학습된 기술이 취약해 새로운 상황에 일반화되지 못하고, 고급 AI 비서가 자신의 윤리적 영향과 관련된 복잡한 개념을 표현하지 못해 과업에 실패하는 리스크.
The risk that a skill is not required during training or is brittle and not generalisable to new situations, and that advanced AI assistants cannot represent complex concepts pertinent to their own ethical impact, so they fail at the task.
① Description
② L3 mapping
③ Duplicate
RAI4-1128
AI 비서 편익·피해 평가 지표 부재
Lack of metrics for evaluating assistant benefits and harms
비서가 초래하는 편익이나 피해의 특정 측면을, 특히 사회의 많은 부분을 포괄할 만큼 광범위한 의미에서 평가할 지표를 개발하기 어려워, 시스템의 피해 위험을 평가하지 못하는 리스크.
The risk that it is challenging to develop metrics for evaluating particular aspects of the benefits or harms caused by an assistant, especially in a sufficiently expansive sense involving much of society, leaving the system's risk of harm unassessed.
① Description
② L3 mapping
③ Duplicate
RAI4-1267
투명성과 설명 가능성
Transparency and explainability
AI 시스템이 대체로 불투명하여 사용자가 그 판단의 근거를 이해할 수 없고, 그 결과 불신과 도입 기피가 생기며 시스템의 행위에 책임을 묻기 어려워지는 리스크.
The risk that AI systems are typically opaque, making it difficult for users to understand the rationale behind their judgements, generating suspicion and reluctance to adopt the technology and making it harder to hold the systems accountable for their actions.
Source members (4)
Source: min_cos=0.7642
RAI4-0610AI 불투명성에 따른 행동 관리 곤란
RAI4-0978의사결정 투명성
RAI4-1226이해관계자 대상 설명가능성 부재
RAI4-1267투명성과 설명 가능성
① Description
② L3 mapping
③ Duplicate
RAI4-1271
견고성과 신뢰성
Robustness and reliability
악의적 공격자나 환경 잡음, 다른 구성요소의 고장으로 입력 데이터가 비정상적으로 변할 때 AI 기반 모델의 성능이 안정적으로 유지되지 못해, 신뢰할 수 없는 모델과 오류에 취약한 에이전트가 되는 리스크.
The risk that an AI-based model's performance is not stable after abnormal changes in the input data caused by a malicious attacker, environmental noise, or a crash of other components, leaving unreliable models and error-prone agents in practice.
① Description
② L3 mapping
③ Duplicate
RAI4-1280
검증 가능성 부족
Lack of verifiability
AI 기반 해법의 비선형적이고 복잡한 구조 탓에 예측과 의사결정의 근거를 알 수 없는 블랙박스가 되어 코드 검증이 이뤄지지 못하고, 의료와 군사 서비스 같은 응용에서 이것이 용납될 수 없는 리스크.
The risk that the non-linear and complex structure of AI-based solutions makes them black boxes providing no information about what drives their predictions and decisions, so the lack of code verification may not be tolerable in applications such as medical healthcare and military services.
① Description
② L3 mapping
③ Duplicate
RAI4-1286
미지 출처에서 획득한 비우호적 AI
Unfriendly AI from an unknown external source
고급 지능형 소프트웨어를 미지의 출처에서 완제품 형태로 획득해야 하는 경우, 예컨대 SETI 연구에서 얻은 신호로부터 추출한 AI가 인간에게 우호적이라는 보장이 없는 리스크.
The risk that advanced intelligent software has to be obtained as a complete package from some unknown source, for example an AI extracted from a signal obtained in SETI research, which is not guaranteed to be human friendly.
① Description
② L3 mapping
③ Duplicate
RAI4-1331
설명가능성 결손에 따른 시정·책임 추적 불능
Explainability deficits impeding rectification and accountability
딥러닝으로 대표되는 AI 알고리즘의 복잡한 내부 동작과 블랙박스 또는 그레이박스 추론으로 산출물이 예측하거나 추적할 수 없게 되어, 이상이 생겼을 때 신속히 시정하거나 책임 소재를 추적할 수 없는 리스크.
The risk that the complex internal workings and black-box or grey-box inference of AI algorithms such as deep learning result in unpredictable and untraceable outputs, making it challenging to quickly rectify them or trace their origins for accountability when anomalies arise.
① Description
② L3 mapping
③ Duplicate
RAI4-1332
모델 강건성 결손
Model robustness deficits
심층 신경망이 비선형적이고 규모가 커서 AI 시스템이 복잡하고 변화하는 운영 환경이나 악의적 간섭과 유도에 취약해져 성능 저하와 의사결정 오류가 발생하는 리스크.
The risk that, as deep neural networks are normally non-linear and large in size, AI systems are susceptible to complex and changing operational environments or malicious interference and inductions, leading to reduced performance and decision-making errors.
① Description
② L3 mapping
③ Duplicate
RAI4-1352
블랙박스 모델 불투명성
Black-box model opacity
수천억 개의 내부 연결을 가진 심층 신경망 기반 생성형 AI의 내부 의사결정 과정이 전문가조차 추적·해석할 수 없게 되어 특정 입력과 출력의 대응을 설명할 수 없는 리스크.
The risk that the internal decision-making processes of generative AI built on deep neural networks with hundreds of billions of internal connections become untraceable and uninterpretable even to the most advanced expert observers, so that developers cannot explain why specific inputs correspond to specific outputs.
① Description
② L3 mapping
③ Duplicate
RAI4-1379
불투명한 상류 구성요소 통합
Opaque upstream component integration
부적절하게 획득되거나 처리·정제되지 않은 데이터를 포함한 상류 제3자 구성요소가 불투명하고 추적 불가능하게 통합되고 수명주기 전반의 공급업체 검증이 미흡하여 하류 사용자에 대한 투명성과 책임성이 저하되는 리스크.
The risk that non-transparent or untraceable integration of upstream third-party components, including data improperly obtained or not processed and cleaned, together with improper supplier vetting across the AI lifecycle, diminishes transparency and accountability for downstream users.
① Description
② L3 mapping
③ Duplicate
RAI4-1390
블랙박스 불신뢰성 사고
Black-box unreliability accidents
범용 AI 모델이 개발자조차 완전히 제어·이해할 수 없는 블랙박스 모델이어서 신뢰성 결여로 예기치 못한 고장이 발생하고, 개발·시험·배포 중 실세계 시스템과 연결될 경우 사고로 이어지는 리스크.
The risk that general purpose AI models, being black-box models not fully controllable or understandable even to their developers, suffer unexpected failures arising from their unreliability, leading to accidents if connected to real-world systems during development, testing, or deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-1409
안전성 미보장 시스템으로 인한 피해
Harm from unsafe underperforming systems
최악 상황 성능이 보장되지 않은 AI 시스템이 주행·의료·전쟁 등 안전필수 영역에 배치되어 연쇄 오류와 사고로 인명 손실, 경제적 피해, 사회 불안이 발생하는 리스크.
The risk that AI systems whose worst-case performance cannot be ensured or proven are implemented in high-stakes, safety-critical domains such as driving, medicine, and warfare, producing cascading errors and accidents that result in loss of life, economic damage, and social unrest.
① Description
② L3 mapping
③ Duplicate
RAI4-1429
투명성과 해석 가능성 부족
Lack of transparency and interpretability
프론티어 AI가 해석하기 어렵고 투명성이 결여되며 학습 데이터에 대한 맥락적 이해가 명시적으로 내장되어 있지 않아, 미세조정이나 인간 피드백 기반 강화학습 없이는 소외 집단의 관점과 수행해야 할 한계를 포착하지 못하는 리스크.
The risk that frontier AI is difficult to interpret and lacks transparency, with contextual understanding of the training data not explicitly embedded, so that without fine tuning or reinforcement learning with human feedback it fails to capture the perspectives of underrepresented groups or the limitations within which it is expected to perform.
① Description
② L3 mapping
③ Duplicate
RAI4-1461
개발 문서화 미흡
Insufficient development documentation
AI 시스템 개발 과정의 결정과 조치가 문서화되지 않아 개발 프로세스 최적화와 시스템의 감사가능성이 훼손되는 리스크.
The risk that decisions and actions taken throughout the development of an AI system go undocumented, undermining optimization of the development process itself and the auditability of the AI system.
① Description
② L3 mapping
③ Duplicate
RAI4-1462
최종 사용자에게 부적절한 투명성 수준
Inappropriate degree of transparency to end users
최종 사용자에 대한 투명성이 설계에 적절히 통합되지 않아 올바른 운용이 저해되고 AI 애플리케이션의 오용이 유발되는 리스크.
The risk that transparency to end users is not adequately integrated into the design of an AI system, preventing its proper operation and causing potential misuse of the AI application.
① Description
② L3 mapping
③ Duplicate
RAI4-1463
하드웨어 연산·전력 요구사항 누락
Omitted compute and power hardware requirements
AI 시스템의 개발과 운영에 필요한 상당한 연산·전력 요구가 하드웨어 선정에서 고려되지 않아 개발과 운영상 문제가 발생하는 리스크.
The risk that the significant computational and power demands of AI system development and operation are not considered in hardware selection, creating issues in development and operation.
① Description
② L3 mapping
③ Duplicate
RAI4-1464
신뢰할 수 없는 데이터 소스 사용
Use of untrustworthy data sources
특히 제3자 데이터 소스를 활용할 때 신뢰할 수 없는 데이터 소스를 선택하여 데이터 품질 요구사항이 충족되지 못하는 리스크.
The risk that choosing an untrustworthy data source, especially when third-party data sources are used to develop the AI system, leaves data quality requirements unfulfilled.
① Description
② L3 mapping
③ Duplicate
RAI4-1465
데이터 이해 부족
Lack of data understanding
사용 데이터에 대한 이해 부족으로 데이터 결함이 간과되고 의도된 기능에 가장 적합한 AI 시스템의 개발이 저해되는 리스크.
The risk that insufficient understanding of the data used for developing an AI system leaves data shortcomings unaddressed and hinders development of a system best suited to the intended functionality.
① Description
② L3 mapping
③ Duplicate
RAI4-1466
잘못된 데이터 라벨
Incorrect data labels
데이터 레이블이 부정확하여 지도학습 AI 시스템이 실측 진실과 의도된 기능을 학습하지 못하는 리스크.
The risk that incorrect data labels prevent a supervised learning AI system from learning the ground truth and therefore the intended functionality.
① Description
② L3 mapping
③ Duplicate
RAI4-1472
과적합 및 과소적합
Over- and underfitting
모델이 훈련 데이터에 과도하게 또는 불충분하게 적응하여 운영 데이터에 직면했을 때 AI 시스템이 신뢰할 수 없게 동작하는 리스크.
The risk that over- or underfitting, the excessive or insufficient adaption of a model to training data, causes an AI system to behave unreliably when confronted with operational data.
① Description
② L3 mapping
③ Duplicate
RAI4-1473
불투명성에 의한 결함 진단 저해
Opacity-impeded model debugging
블랙박스 모델 기반 AI 시스템의 설명가능성이 제한되어 개발자가 데이터나 모델 자체의 결함을 탐지하지 못하고 시스템의 성능과 안전 수준이 저하되는 리스크.
The risk that the limited explainability of AI systems based on black-box models prevents developers from detecting shortcomings in the data or the model itself, decreasing the performance and safety levels of the AI system.
① Description
② L3 mapping
③ Duplicate
RAI4-1474
코너 케이스의 신뢰성 없음
Unreliability in corner cases
AI 시스템이 희귀하거나 모호한 입력 데이터인 코너 케이스에 직면할 때 통제된 동작이 요구됨에도 신뢰할 수 없는 동작을 보이는 리스크.
The risk that an AI system shows unreliable behavior when confronted with rare or ambiguous input data, also called corner cases, where controlled behavior is required.
① Description
② L3 mapping
③ Duplicate
RAI4-1475
신뢰도 추정 기능 결여
Missing or faulty confidence estimation
AI 시스템이 출력에 상응하는 신뢰도 수준을 제공하지 못하거나 잘못 제공하여 성능과 안전에 부정적 영향이 발생하는 리스크.
The risk that an AI system fails to provide, or incorrectly provides, a level of confidence corresponding to its output, negatively impacting performance and safety.
Source members (2)
Source: min_cos=0.8331 · Mixed L3
RAI4-1475신뢰도 추정 기능 결여
RAI4-1717타당성·신뢰성 부족에 의한 부정확한 출력
① Description
② L3 mapping
③ Duplicate
RAI4-1476
운영 데이터 분포 편차
Operational data distribution deviation
테스트 세트가 근사한 분포와 실제 운영 데이터 분포 사이의 예기치 못한 편차로 배포된 AI 애플리케이션이 신뢰할 수 없게 동작하는 리스크.
The risk that an unexpected deviation between the distribution approximated by the test set and the actual operational data causes a deployed AI application to behave unreliably.
① Description
② L3 mapping
③ Duplicate
RAI4-1478
컨셉 드리프트
Concept drift
입력 변수와 모델 출력 간의 관계가 변화하는 개념 드리프트가 적절히 처리되지 않아 AI 시스템의 신뢰성이 저하되는 리스크.
The risk that concept drift, a change in the relationship between input variables and model output, is not treated appropriately and reduces the reliability of AI systems.
① Description
② L3 mapping
③ Duplicate
RAI4-1485
AI 구성요소 상호작용의 원인 규명 곤란
Unclear attribution of harm from AI component interactions
서로 다른 AI 구성요소 간의 상호작용이 피해를 유발하지만 어떤 구성요소가 원인인지 특정하기 어려운 리스크.
The risk that interactions between different AI components cause harm while it remains difficult to pinpoint which components are the cause.
① Description
② L3 mapping
③ Duplicate
RAI4-1495
과도한 안전 튜닝
Overly restrictive safety tuning
과도한 안전 훈련이나 안전 튜닝이 AI 시스템의 성능을 저하시켜 지나치게 조심스러운 행동을 유발하고, 유해한 프롬프트와 부분적으로 유사한 완전히 안전한 프롬프트에 대해서도 응답을 거부하게 되는 리스크.
The risk that excessive safety training or safety tuning impairs the performance of AI systems, leading to overly cautious behavior in which they refuse to answer entirely safe prompts that are partially similar to harmful ones.
① Description
② L3 mapping
③ Duplicate
RAI4-1500
역량 식별·측정 곤란
Capability identification and measurement difficulty
범용 AI 시스템의 역량이 잠재적 위험의 분포가 넓고 이를 평가할 명확한 지표가 없으며 예측 불가능한 창발적 속성이 존재하여 고정 목적 AI에 비해 측정하기 어려운 리스크.
The risk that the capabilities of general-purpose AI systems are difficult to measure compared with those of more limited and fixed-purpose AI systems, due in part to a broader distribution of potential risks, a lack of well-defined metrics to evaluate them, and risks from unpredictable or emergent model properties.
① Description
② L3 mapping
③ Duplicate
RAI4-1502
내재 가치 측정 부정확
Inaccurate measurement of encoded values
AI 시스템의 출력이 인간의 가치에 확고히 부합하는지 아니면 부분적으로만 상관된 모방인지 평가할 강건한 프레임워크가 부재하고, 모델이 학습한 가치 표상이 출력에 온전히 반영되지 않으며 훈련·배포 단계에 따라 어떻게 변하는지 알려지지 않아 내재 가치의 측정이 부정확해지는 리스크.
The risk that, lacking robust frameworks for evaluating whether AI outputs robustly conform to human values rather than merely mimicking them, and with outputs imperfectly reflecting the learned value representations whose evolution across training and deployment stages is unknown, measurement of encoded values becomes inaccurate.
① Description
② L3 mapping
③ Duplicate
RAI4-1504
인간 평가 한계 초과 출력
Outputs beyond human evaluability
인간 피드백을 이용한 평가로 AI 모델을 훈련할 때 평가자가 감지하기 어려운 오류를 포함한 출력을 정답과 유사하게 긍정 평가하여, 모델이 소프트웨어 취약점이 있는 코드나 정치적으로 편향된 정보처럼 미묘하게 잘못되거나 유해한 출력을 학습하고 극단적으로는 숨겨진 오류나 백도어를 포함한 출력을 생성하는 리스크.
The risk that, when AI models are trained through evaluation with human feedback, human evaluators rate outputs containing hard-to-detect errors positively or similarly to correct ones, so the model learns to produce subtly incorrect or harmful outputs such as code with software vulnerabilities or politically biased information, and in extreme cases complicated outputs containing hidden errors or backdoors.
① Description
② L3 mapping
③ Duplicate
RAI4-1543
아웃바운드 통신에 의한 기밀 유출·무단 행위
Data leakage and unauthorized actions from unintended outbound communication
네트워크 접근 권한을 폭넓게 가진 AI 시스템이 통신 채널 화이트리스트나 최소권한 원칙 없이 배포되어, 제공자·배포자·사용자가 의도하지 않은 아웃바운드 통신으로 기밀 데이터가 유출되거나 원치 않는 행위가 실행되는 리스크.
The risk that an AI system with broad network access, deployed without communication whitelisting or least-privilege constraints, sends data outbound in ways no provider, deployer, or user intended, leaking confidential data or performing unwanted actions.
① Description
② L3 mapping
③ Duplicate
RAI4-1551
센서 드리프트에 의한 배포 시스템 성능 저하
Deployed-system degradation from sensor and distribution drift
물리 센서와 데이터 소스에 의존하는 배포된 AI 시스템에서 하드웨어 드리프트에 따른 데이터 분포 변화가 발생하여 시스템의 견고성과 성능이 저하되는 리스크.
The risk that hardware drift in the physical sensors and data sources of deployed AI systems causes data distribution drift that degrades system robustness and performance.
① Description
② L3 mapping
③ Duplicate
RAI4-1633
인프라 제어 오판단에 의한 필수 서비스 붕괴
Essential-service collapse from erroneous AI infrastructure control
전력망·수처리·통신·교통 조정 시스템에 배포된 범용 AI가 운영 데이터를 오해석하거나 연쇄 장애를 예측하지 못한 제어 결정을 내려 상호연결된 인프라가 불안정해지고 광범위한 정전, 수질 오염, 통신 두절 등 필수 서비스가 붕괴되는 리스크.
The risk that general-purpose AI deployed in power grid, water treatment, telecommunications, or transportation systems misinterprets operational data, fails to anticipate cascading failure modes, or makes control decisions that destabilize interconnected infrastructure, causing widespread blackouts, contaminated water supplies, communications breakdowns, and collapse of essential services.
① Description
② L3 mapping
③ Duplicate
RAI4-1703
AI 시스템 요구사항·목적 명세 오류
Misspecification of AI system requirements and purpose
AI 시스템의 요구사항과 목적이 잘못 이해되거나 잘못 도출되어 부정확한 요구사항 도출, 차선의 모델링·하이퍼파라미터 선택, 변화하는 도메인 요구에 대한 적응 실패가 발생하는 리스크.
The risk that the requirements and purpose of an AI system are misunderstood or mis-elicited, resulting in incorrect requirement elicitation, suboptimal modelling or hyperparameter choices, and failure to adapt to changing domain requirements.
① Description
② L3 mapping
③ Duplicate
RAI4-1705
모니터링 부족에 의한 미탐지 안전·프라이버시 위반
Undetected safety and privacy violations from insufficient monitoring
배포된 AI에 대한 모니터링과 해석 가능성이 부족하여 블랙박스 불투명성이 인간의 주체성을 축소하고 윤리·안전 원칙 위반과 프라이버시 침해가 탐지되지 않은 채 남는 리스크.
The risk that insufficient monitoring and interpretability of deployed AI leaves black-box opacity diminishing human agency and allows ethical or safety-principle violations and privacy violations to go undetected.
① Description
② L3 mapping
③ Duplicate
RAI4-1719
보안·복원력 부족에 의한 공격 취약과 복구 실패
Attack susceptibility and recovery failure from lack of security and resilience
AI 시스템이 적대적 공격, 데이터 오염, 정보 유출에 취약하거나 악영향을 견디고 회복하지 못하는 리스크.
The risk that an AI system is susceptible to adversarial attacks, data poisoning, or exfiltration, or is unable to withstand and recover from adverse events.
① Description
② L3 mapping
③ Duplicate
RAI4-1721
설명 가능성 및 해석 가능성 부족
Lack of explainability and interpretability
AI 시스템 작동의 기저 메커니즘을 표현하거나 출력의 의미를 맥락에 맞게 전달하지 못하는 리스크.
The risk of inability to represent the mechanisms underlying an AI system's operation or to convey the meaning of its outputs in context.
① Description
② L3 mapping
③ Duplicate

상호작용 안전성 · Interaction Safety · 219 cards

RAI3-G-INT-01 폭력 Violence10 cards
타인·집단·동물에 대한 물리적·정신적 해를 가하거나, 그 위협·조장·미화를 포함하는 콘텐츠
IDCardHuman audit
RAI4-0426
혐오발언 생성
Hate speech generation
생성 시스템이 특정 정체성 집단을 표적으로 하는 혐오적·모욕적 콘텐츠를 산출하는 리스크.
The risk that generative systems produce hateful or abusive content targeting identity groups.
① Description
② L3 mapping
③ Duplicate
RAI4-0674
문화적 박탈
Cultural dispossession
말하기 방식, 유머 표현, 문화 정체성을 구성하는 소리와 목소리 등 문화적 재화와 가치가 의도적 또는 비의도적으로 소거되거나 다른 문화에서 부적절하게 재사용되는 리스크
The risk of intentional or unintentional erasure of cultural goods and values, such as ways of speaking, expressing humour, or sounds and voices that contribute to a cultural identity, or their inappropriate re-use in other cultures.
① Description
② L3 mapping
③ Duplicate
RAI4-0678
신체적 위해를 유발하는 출력
Model outputs leading to physical harm
모델이 명백히 폭력적이거나 은밀하게 위험하거나 그 밖에 간접적으로 안전하지 않은 언어를 생성하여 신체적 위해로 이어지는 리스크
The risk that a model generates overtly violent, covertly dangerous, or otherwise indirectly unsafe language that leads to physical harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0705
폭력 조장 위험 콘텐츠
Violence-inciting dangerous content
폭력적·선동적·급진화·위협적 콘텐츠의 제작과 접근이 용이해지고 자해나 불법 활동의 수행을 권장하는 콘텐츠가 산출되며, 증오·비하·고정관념 콘텐츠에 대한 대중의 노출을 통제하기 어려워지는 리스크
The risk of eased production of and access to violent, inciting, radicalizing, or threatening content and recommendations to carry out self-harm or conduct illegal activities, including difficulty controlling public exposure to hateful, disparaging, or stereotyping content.
① Description
② L3 mapping
③ Duplicate
RAI4-0727
개인·집단·조직에 대한 언어적 공격
Verbal attacks on individuals, groups, or organizations
챗봇이 개인, 집단 또는 조직을 언어적으로 공격하거나 훼손하는 리스크
The risk that a chatbot verbally attacks or undermines an individual, group, or organization.
① Description
② L3 mapping
③ Duplicate
RAI4-1049
조직 재정·평판 손상
Organizational financial and reputational damage
ML 시스템을 구축하거나 사용하는 조직의 재정적 및/또는 평판 손상 위험.
The risk of financial and/or reputational damage to the organization building or using the ML system.
① Description
② L3 mapping
③ Duplicate
RAI4-1070
유독한 언어
Toxic language
LM이 욕설, 정체성 공격, 모욕, 위협, 성적으로 노골적인 내용, 비하 표현, 폭력 선동 등 증오 표현이나 유독한 언어를 예측·생성하는 리스크.
The risk that LMs predict hate speech or other toxic language, including profanities, identity attacks, insults, threats, sexually explicit content, demeaning language, or language inciting violence targeted at a person or group because of innate characteristics.
Source members (4)
Source: min_cos=0.7697 · Mixed L3
RAI4-0726학습 데이터 내 독성 언어
RAI4-0729정체성 공격형 유해 언어
RAI4-1052증오심 표현 및 공격적인 언어
RAI4-1070유독한 언어
① Description
② L3 mapping
③ Duplicate
RAI4-1177
공격적 콘텐츠 생성
Offensive content generation
LLM이 위협과 모욕, 경멸, 욕설, 빈정거림, 무례함 같은 공격적 콘텐츠나 행위를 식별하고 반대하지 못한 채 그러한 콘텐츠를 생성하는 리스크.
The risk that LLMs fail to identify and oppose offensive content or actions involving threat, insult, scorn, profanity, sarcasm, and impoliteness, and generate such content instead.
Source members (4)
Source: min_cos=0.7998
RAI4-0636LLM 악용에 의한 독성 콘텐츠 생성
RAI4-0690LLM의 명시적·암묵적 독성 콘텐츠 생성
RAI4-1177공격적 콘텐츠 생성
RAI4-1188폭력적 콘텐츠 생성
① Description
② L3 mapping
③ Duplicate
RAI4-1308
독성 텍스트 생성
Toxic text generation
LLM이 프롬프트를 받았을 때 혐오 발언과 모욕적 언어, 폭력적 발언, 욕설을 아우르는 독성 텍스트를 생성하는 리스크.
The risk that an LLM generates toxic text when prompted, toxicity being an umbrella term encompassing hate speech, abusive language, violent speech, and profane language.
① Description
② L3 mapping
③ Duplicate
RAI4-1447
폭력/무력 충돌
Violence/armed conflict
기술 시스템을 사용하거나 오용하여 사이버 공격, 보안 침해, 치명적 생화학 무기 개발을 선동·촉진·수행함으로써 폭력과 무력 충돌이 초래되는 리스크.
The risk that use or misuse of a technology system to incite, facilitate, or conduct cyberattacks, security breaches, and lethal biological and chemical weapons development results in violence and armed conflict.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-02 성적 콘텐츠 Sexual17 cards
성적 행위·성범죄·성착취·아동 성적 콘텐츠·성상품화 등 성 관련 유해 콘텐츠
IDCardHuman audit
RAI4-0627
허위 콘텐츠에 의한 개인 표적 피해
Targeted harm to individuals via fake content
악의적 행위자가 범용 AI로 허위 콘텐츠를 생성하여 사기, 갈취, 심리적 조작, 비동의 성적 이미지와 아동 성착취물 제작, 개인·조직에 대한 표적 방해에 사용함으로써 개인에게 표적 피해를 입히는 리스크
The risk that malicious actors use general-purpose AI to generate fake content that harms individuals in a targeted way, including scams, extortion, psychological manipulation, non-consensual intimate imagery and child sexual abuse material, and targeted sabotage of individuals and organisations.
① Description
② L3 mapping
③ Duplicate
RAI4-0641
비합의 성적 대상화
Non-consensual sexualization
기술이나 애플리케이션을 이용해 개인 또는 집단을 동의 없이 성적으로 대상화하는 리스크
The risk of the non-consensual sexualisation of an individual or group using a technology or application.
① Description
② L3 mapping
③ Duplicate
RAI4-0642
증오·모욕·외설 콘텐츠 생성
Intentional generation of hateful and obscene content
생성형 AI 모델이 증오·모욕·모독(HAP) 또는 외설적 콘텐츠를 생성하는 데 의도적으로 사용되는 리스크
The risk that generative AI models are used intentionally to generate hateful, abusive, and profane (HAP) or obscene content.
Source members (2)
Source: min_cos=0.8589 · Mixed L3
RAI4-0642증오·모욕·외설 콘텐츠 생성
RAI4-0689증오·모욕·외설 출력
① Description
② L3 mapping
③ Duplicate
RAI4-0670
집단 왜곡 표상과 독성 콘텐츠 생성
Group misrepresentation and toxic content generation
AI 시스템이 특정 집단을 과소·과대 표현하거나 허위로 표현하고 독성·모욕·학대·증오 콘텐츠를 생성하는 리스크
The risk that AI systems under-, over-, or misrepresent certain groups or generate toxic, offensive, abusive, or hateful content.
① Description
② L3 mapping
③ Duplicate
RAI4-0684
성범죄 조장 콘텐츠 생성
Generation of content enabling sex-related crimes
AI가 성매매 인신매매, 성폭력, 성희롱, 비합의 친밀 콘텐츠 유포, 수간 등 성 관련 범죄의 실행을 가능하게 하거나 조장 또는 승인하는 응답을 산출하는 리스크
The risk that an AI system produces responses that enable, encourage, or endorse the commission of sex-related crimes such as sex trafficking, sexual assault, sexual harassment, nonconsensual sharing of sexually intimate content, and bestiality.
Source members (6)
Source: min_cos=0.7477 · Mixed L3
RAI4-0665AI 기반 아동 성착취물 생성
RAI4-0673아동 성착취 조장 콘텐츠 생성
RAI4-0682비폭력 범죄 조장 콘텐츠 생성
RAI4-0684성범죄 조장 콘텐츠 생성
RAI4-0685음란물 및 성적 대화 콘텐츠 생성
RAI4-0691폭력 범죄 조장 콘텐츠 생성
① Description
② L3 mapping
③ Duplicate
RAI4-0688
커뮤니티 기준 위반 독성 콘텐츠 생성
Generation of community-standard-violating toxic content
유혈, 아동 성적 묘사, 욕설, 정체성 공격 등 커뮤니티 기준을 위반하는 콘텐츠가 생성되어 특정 집단에 피해를 주거나 그들에 대한 증오와 폭력을 선동하는 리스크
The risk of generating content that violates community standards, including harming or inciting hatred or violence against groups, such as gore, sexual content of children, profanities, and identity attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-0714
아동·청소년 유해 콘텐츠 제공
Provision of content harmful to children and youth
LLM이 아동과 청소년에게 유해한 콘텐츠를 담은 답변을 산출하도록 유도될 수 있는 리스크
The risk that LLMs are leveraged to solicit answers that contain content harmful to children and youth.
① Description
② L3 mapping
③ Duplicate
RAI4-0718
외설·비하·학대 이미지 생성 및 접근 용이화
Eased production of and access to abusive imagery
해를 끼칠 수 있는 외설적·굴욕적·모욕적 이미지, 특히 합성 아동 성적 학대 자료(CSAM)와 성인의 비동의 친밀 이미지(NCII)의 제작과 접근이 용이해지는 리스크
The risk that production of and access to obscene, degrading, or abusive imagery which can cause harm is eased, including synthetic child sexual abuse material (CSAM) and nonconsensual intimate images (NCII) of adults.
① Description
② L3 mapping
③ Duplicate
RAI4-0728
독성 콘텐츠 생성
Generation of toxic content
모델이 무례하고 모욕적이며 심지어 불법적인 정보를 담은 독성 콘텐츠를 생성하는 리스크
The risk that a model generates toxic content containing rude, disrespectful, and even illegal information.
Source members (7)
Source: min_cos=0.7575 · Mixed L3
RAI4-0710유해·차별적 콘텐츠 생성
RAI4-0712비윤리적·유해 콘텐츠 및 고위험 조언 생성
RAI4-0728독성 콘텐츠 생성
RAI4-0936독성 및 악의적인 콘텐츠
RAI4-1165모욕적 콘텐츠 생성
RAI4-1220유해·부적절 콘텐츠 생성
RAI4-1349의도치 않은 유해 콘텐츠 생성
① Description
② L3 mapping
③ Duplicate
RAI4-0828
참여 극대화 콘텐츠에 의한 자율성 훼손
Erosion of user autonomy from engagement-maximizing content
생성형 AI가 진실성과 무관하게 클릭베이트 헤드라인과 대량의 기사·웹페이지를 생성하여 검색 노출과 클릭을 극대화함으로써, 이용자의 탐색 행동을 조작하고 이용 경험과 소비자 자율성을 훼손하는 리스크
The risk that generative AI produces clickbait headlines and mass articles regardless of veracity to maximize search visibility and clicks, manipulating how users navigate the internet and applications, degrading their experience, and undermining consumer autonomy.
① Description
② L3 mapping
③ Duplicate
RAI4-1008
기술을 이용한 폭력
Technology-facilitated violence
알고리즘 기능이 시스템을 괴롭힘과 폭력에 이용할 수 있게 하여 생성 AI의 비합의 성적 이미지 생성, 신상털기, 트롤링, 사이버스토킹·괴롭힘, 감시·통제 등 온라인 폭력이 발생하는 리스크.
The risk that algorithmic features enable use of a system for harassment and violence, including non-consensual sexual imagery in generative AI, doxxing, trolling, cyberstalking, cyberbullying, monitoring and control, and online harassment and intimidation.
① Description
② L3 mapping
③ Duplicate
RAI4-1136
대규모 유해 콘텐츠 생성
Scaled harmful content generation
적절한 안전·보안 장치 없이 고급 AI 비서가 위협 행위자로 하여금 아동 성학대 자료, 사기, 허위정보 같은 유해 콘텐츠를 더 빠르고 정확하며 저렴하고 개인화된 형태로 더 넓은 범위에 생성하게 하는 리스크.
The risk that, without proper safety and security mechanisms, advanced AI assistants allow threat actors to create harmful content such as child sexual abuse material, fraud, and disinformation more quickly, accurately, cheaply, and with greater personalization and reach.
① Description
② L3 mapping
③ Duplicate
RAI4-1137
합의되지 않은 콘텐츠 대규모 생성
Non-consensual content generation at scale
생성 AI가 나체, 혐오, 폭력 묘사와 개인의 초상을 이용한 합의되지 않은 콘텐츠를 만들고, 비서의 도구 사용과 계획 능력이 개인 표적화와 착취·괴롭힘·협박을 자동화하고 대규모로 확산시키는 리스크.
The risk that generative AI produces non-consensual content depicting nudity, hate, or violence and using individuals' likenesses, while assistants' tool-use and planning capabilities automate targeting of individuals for exploitation, harassment, and blackmail at scale.
① Description
② L3 mapping
③ Duplicate
RAI4-1204
비합의 성적 딥페이크
Non-consensual sexual deepfakes
생성 AI가 합의되지 않은 성적 이미지와 영상의 딥페이크를 만들어 타인을 해치거나 굴욕감을 주거나 성적으로 대상화하는 데 악의적으로 사용되는 리스크.
The risk that generative AI is maliciously used to harm, humiliate, or sexualize another person by generating deepfakes of nonconsensual sexual imagery or videos.
Source members (3)
Source: min_cos=0.7972 · Mixed L3
RAI4-0638동의 없는 딥페이크 인물 모사
RAI4-1204비합의 성적 딥페이크
RAI4-1357성적 노골 콘텐츠 오용
① Description
② L3 mapping
③ Duplicate
RAI4-1542
외부 도구 연동을 통한 유해 콘텐츠 유입
Ingestion of harmful content through external tool integration
외부 도구·플러그인과의 통합과 상호 연결이 확대됨에 따라 악의적 외부 입력에 노출되어 유해 콘텐츠가 시스템에 유입되는 리스크.
The risk that growing integration and interconnectivity with external tools and plugins exposes a system to malicious external inputs that introduce harmful content.
① Description
② L3 mapping
③ Duplicate
RAI4-1555
취약점 기반 맞춤형 괴롭힘 콘텐츠 생성
Vulnerability-targeted personalized harassment content generation
GPAI가 표적 개인의 약점에 맞춘 콘텐츠를 자동 생성하는 데 오용되어 괴롭힘·갈취·협박의 효율성과 성공률이 높아지는 리스크.
The risk that GPAI is misused to automatically generate content personalized to targets' weak spots, making harassment, extortion, and intimidation more efficient and more likely to succeed.
① Description
② L3 mapping
③ Duplicate
RAI4-1582
비동의 성적 이미지 생성
Non-consensual intimate imagery generation
성인의 실제 외모를 이용한 노골적 성적 자료가 본인의 동의 없이 생성되는 리스크.
The risk that sexually explicit material is created using an adult person's likeness without their consent.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-03 자해 Self-harm5 cards
자살·자해·위험 약물 남용·극단적 다이어트 등 개인의 신체·정신 안전을 직접 위협하는 콘텐츠
IDCardHuman audit
RAI4-0023
에이전트의 자해 조장
Self-harm facilitation by agents
에이전트가 자해 또는 대인 위해의 가능성을 높이는 행동 지향적 지원·계획·자원을 제공하는 리스크.
The risk that an agent provides action-oriented support, planning, or resources that increase the likelihood of self-harm or interpersonal harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0157
챗봇을 통한 자해 격려
Self-harm encouragement by chatbot
챗봇이 자해 사고나 행동을 강화하거나 정상화하거나 중단시키지 못하는 리스크.
The risk that a chatbot reinforces, normalizes, or fails to interrupt self-harm ideation or behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0687
자살·자해 조장 콘텐츠 생성
Generation of content encouraging suicide and self-harm
AI가 자살, 자해, 섭식장애 등 의도적 자해 행위를 가능하게 하거나 조장 또는 승인하는 응답을 산출하는 리스크
The risk that an AI system produces responses that enable, encourage, or endorse acts of intentional self-harm such as suicide, self-injury, and disordered eating.
① Description
② L3 mapping
③ Duplicate
RAI4-0833
정신건강 유해 콘텐츠
Mental health harmful content
모델이 자살을 조장하거나 공황·불안을 유발하는 등 정신건강에 위험한 응답을 생성하여 이용자의 정신건강에 부정적 영향을 미치는 리스크
The risk that a model generates risky responses about mental health, such as content that encourages suicide or causes panic or anxiety, negatively affecting users' mental health.
Source members (2)
Source: min_cos=0.8349
RAI4-0833정신건강 유해 콘텐츠
RAI4-0839신체 건강 위해 콘텐츠
① Description
② L3 mapping
③ Duplicate
RAI4-0861
자해 촉진
Self-harm facilitation
기술 시스템 사용의 직접적 또는 간접적 결과로 이용자가 자신의 신체에 고의적 손상을 가하게 되는 리스크
The risk that a person deliberately damages their own body as a direct or indirect result of using a technology system.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-04 혐오·차별 Hate and Unfairness9 cards
인종, 지역, 국적, 민족, 가정형태, 공공, 성, 세대/나이, 신체 조건, 유명인, 출산 및 혼인 여부, 장애/병력, 재난 및 범죄 피해, 종교/신념, 직업/학력/사회적 지위, 취미 및 욕설/비속어등 + 종교
IDCardHuman audit
RAI4-0411
언어별 피해 과소탐지
Language-specific harm under-detection
평가 범위가 취약한 언어와 방언에서 유해·모욕적·비안전 콘텐츠가 덜 안정적으로 탐지되는 리스크.
The risk that toxic, disrespectful, or unsafe content is less reliably detected in languages and dialects with weaker evaluation coverage.
① Description
② L3 mapping
③ Duplicate
RAI4-0683
출력 단계 집단 표상 편향
Group misrepresentation bias in model outputs
생성된 콘텐츠가 특정 집단이나 개인을 불공정하게 표상하는 리스크
The risk that generated content unfairly represents certain groups or individuals.
① Description
② L3 mapping
③ Duplicate
RAI4-0706
보호 속성 기반 부당 대우
Protected-attribute unfair treatment
인종, 민족, 연령, 성별, 성적 지향, 종교, 출신 국가, 혼인 여부, 장애, 언어 등 보호 속성을 근거로 개인이 불공정하거나 부적절한 대우 또는 자의적 차별을 받는 리스크
The risk of unfair or inadequate treatment or arbitrary distinction based on a person's race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other protected groups.
① Description
② L3 mapping
③ Duplicate
RAI4-0720
편견과 과소표현으로 인한 위험
Risks from bias and underrepresentation
영어권·서구 편중 학습 데이터로 인해 인종, 성별, 문화, 연령, 장애 전반에 편향된 출력이 산출되어 의료, 채용, 대출 등 고위험 영역에서 과소대표 사용자에게 피해를 주는 리스크
Training corpora that overrepresent English-speaking and Western populations produce outputs biased across race, gender, culture, age, and disability, harming underrepresented users in high-stakes domains such as healthcare, recruitment, and lending.
① Description
② L3 mapping
③ Duplicate
RAI4-0937
사회 집단에 대한 부당한 부정적 편견
Unfairly negative bias against social groups
모델이 일방적이거나 부정확한 정보에 기반해 성별·인종·종교 등에 대한 부정적 고정관념과 결부된, 사회 집단이나 개인에 대한 부당하게 부정적인 태도를 드러내는 리스크.
The risk that a system exhibits an unfairly negative attitude towards a social group or individuals based on one-sided or inaccurate information, typically tied to widely disseminated negative stereotypes regarding gender, race, or religion.
① Description
② L3 mapping
③ Duplicate
RAI4-1068
고정관념 인코딩에 의한 차별적 대우
Discriminatory treatment from encoded stereotypes
차별적 언어와 사회적 고정관념을 인코딩한 언어모델이 성별·종교·성적지향·장애·연령 등 민감한 속성에 따른 차별적 대우와 자원 접근 격차 등 다양한 피해를 야기하는 리스크.
The risk that language models encoding discriminatory language or social stereotypes cause harms including differential treatment or access to resources based on sensitive traits such as sex, religion, gender, sexual orientation, ability, and age.
① Description
② L3 mapping
③ Duplicate
RAI4-1166
불공정·차별적 산출
Unfair and discriminatory output
모델이 인종과 성별, 종교, 외모 등에 근거한 사회적 편향을 비롯한 불공정하고 차별적인 산출을 생성하여 특정 집단을 불편하게 하고 사회의 안정과 평화를 훼손하는 리스크.
The risk that a model produces unfair and discriminatory data, such as social bias based on race, gender, religion, or appearance, which may discomfort certain groups and undermine social stability and peace.
① Description
② L3 mapping
③ Duplicate
RAI4-1168
민감 주제 편향 콘텐츠
Biased content on sensitive topics
정치를 비롯한 민감하고 논쟁적인 주제에서 언어모델이 특정 정치적 입장을 지지하는 편향되고 오도하며 부정확한 콘텐츠를 생성하여 다른 관점을 차별하거나 배제하는 리스크.
The risk that on some sensitive and controversial topics, especially politics, language models generate biased, misleading, and inaccurate content that supports a specific position and leads to discrimination or exclusion of other viewpoints.
① Description
② L3 mapping
③ Duplicate
RAI4-1174
안전하지 않은 지시 주제
Unsafe instruction topic
입력 지시 자체가 부적절하거나 불합리한 주제를 참조할 때 모델이 그 지시를 따라 광신과 인종주의 같은 안전하지 않은 콘텐츠를 생성하여 사회에 부정적 영향을 줄 수 있는 리스크.
The risk that when the input instructions themselves refer to inappropriate or unreasonable topics, the model follows them and produces unsafe content such as fanaticism or racism with possible negative impact on society.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-05 정치적 중립성 Political Neutrality6 cards
정치적 극단주의 선동, 정치인 비방 및 명예훼손, 선거 절차 방해와 선동, 종교 전복 및 관련 사기 수법, 정치적 동기에 의한 폭력의 정당화,
IDCardHuman audit
RAI4-0403
도덕적 프레이밍 편향
Moral framing bias
프롬프트·응답·평가의 구성 방식이 이용자를 논쟁적 사안에 대한 특정한 도덕적 해석으로 유도하는 리스크.
The risk that the framing of prompts, responses, or evaluations channels users toward a particular moral interpretation of a contested issue.
① Description
② L3 mapping
③ Duplicate
RAI4-0422
도덕적 세계관 배제
Moral-worldview exclusion
종교적·영적·토착적·비세속적 도덕 세계관이 모델 행동과 평가 기준에서 배제되는 리스크.
The risk that religious, spiritual, indigenous, or non-secular moral worldviews are excluded from model behavior and evaluation standards.
① Description
② L3 mapping
③ Duplicate
RAI4-0529
정치 불안정 유발 오용
Politically destabilizing misuse
기술 시스템의 사용 또는 오용이 직간접적으로 정치적 불안을 야기하는 리스크
The risk that the use or misuse of a technology system directly or indirectly causes political unrest.
Source members (2)
Source: min_cos=0.8352
RAI4-0529정치 불안정 유발 오용
RAI4-1450구조적 정치 불안정화
① Description
② L3 mapping
③ Duplicate
RAI4-0939
모델의 극단적·편향적 견해 표출
Expression of extremist and politically biased views
대형 모델이 정치적 주제에서 부적절하거나 극단주의적인 견해를 표출하고, 중립을 표방하면서도 특정 정치 성향의 편향을 드러내는 리스크.
The risk that large models express inappropriate or extremist views on political topics and, while claiming neutrality, exhibit notable political biases across policy domains.
① Description
② L3 mapping
③ Duplicate
RAI4-1159
정치적 영향력 전략 역량
Political influence strategy capability
모델이 행위자가 정치적 영향력을 획득하고 행사하는 데 필요한 사회적 모델링과 계획을, 다수 행위자와 풍부한 사회적 맥락이 있는 시나리오에서까지 수행하는 리스크.
The risk that a model performs the social modelling and planning necessary for an actor to gain and exercise political influence, not just at a micro level but in scenarios with multiple actors and rich social context.
① Description
② L3 mapping
③ Duplicate
RAI4-1451
정치적 조작
Political manipulation
개인 데이터를 사용하거나 오용하여 마이크로 광고나 딥페이크·합성 미디어를 통해 개인의 관심사·성격·취약성을 표적으로 맞춤형 정치 메시지를 전달하는 리스크.
The risk that personal data is used or misused to target individuals' interests, personalities, and vulnerabilities with tailored political messages via micro-advertising or deepfakes and synthetic media.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-06 개인정보 Privacy39 cards
개인식별정보 노출, 의료/금융/위치/생체 정보 노출, 통신내용 침해, 개인사 노출, 개인 신상 정보 수집 방법, 프라이버시 침해 기술
IDCardHuman audit
RAI4-0433
생체정보 기반 프라이버시 침해
Biometric privacy intrusion
얼굴·음성·보행 등 생체인식 AI 시스템이 침해적인 식별과 추적을 가능하게 하는 리스크.
The risk that face, voice, gait, or other biometric AI systems enable intrusive identification and tracking.
① Description
② L3 mapping
③ Duplicate
RAI4-0631
신원 사칭 및 도용
Impersonation and identity theft
제3자가 개인, 집단, 조직의 신원을 도용하여 사기, 조롱, 그 밖의 가해를 함으로써 당사자나 다른 당사자에게 피해가 발생하는 리스크
The risk that a third party steals the identity of an individual, group, or organisation in order to defraud, mock, or otherwise harm them or another party.
① Description
② L3 mapping
③ Duplicate
RAI4-0652
개인화 정보 악용 표적 사기
Targeted fraud exploiting personalized information
생성 모델이 개인화된 정보를 이용해 개별 사용자를 효율적으로 표적화하는 데 오용되어, 피해자의 신뢰를 악용한 설득력 높은 자동화 사기로 민감 정보가 탈취되고 기만의 성공 가능성이 높아지는 리스크
The risk that generative models are misused to target individual users more efficiently using personalized information, producing highly convincing automated fraudulent schemes that exploit victims' trust and extract sensitive data.
① Description
② L3 mapping
③ Duplicate
RAI4-0655
초상 도용
Appropriation of personal likeness
개인의 초상이나 그 밖의 식별 가능한 특징이 사용되거나 변형되는 리스크
The risk that a person's likeness or other identifying features are used or altered.
① Description
② L3 mapping
③ Duplicate
RAI4-0734
기밀정보 무단 공유에 따른 사업 손실
Business loss from unauthorised sharing of confidential information
기업 전략, 재무 계획 등 민감·기밀 정보와 문서가 제3자와 무단으로 공유되어 시장 지위나 수익을 상실하는 리스크
The risk that sensitive, confidential information and documents such as corporate strategy and financial plans are shared with third parties without authorisation, risking loss of market position or revenue.
① Description
② L3 mapping
③ Duplicate
RAI4-0735
개인 데이터의 공개 및 부적절한 공유
Disclosure and improper sharing of personal data
개인의 데이터가 공개되거나 부적절하게 공유되고, AI가 원시 데이터에 명시적으로 담기지 않은 추가 정보를 추론하거나 모델 학습을 위해 개인 데이터를 공유함으로써 새로운 유형의 공개 위험이 생기고 악화되는 리스크
The risk that data of individuals is revealed and improperly shared, with AI creating new types of disclosure risk by inferring additional information beyond what is explicitly captured in the raw data and exacerbating disclosure risks through sharing personal data to train models.
① Description
② L3 mapping
③ Duplicate
RAI4-0739
민감한 정보 노출
Sensitive-information exposure
은폐하도록 사회화된 매우 사적인 민감 정보가 드러나고, 생성 기술이 검열·삭제된 콘텐츠를 재구성하거나 추론된 민감 데이터·선호·의도를 노출함으로써 새로운 유형의 노출 위험이 발생하는 리스크
The risk that sensitive private information people view as deeply primordial and have been socialized into concealing is revealed, with AI creating new types of exposure risk through generative techniques that reconstruct censored or redacted content and through exposing inferred sensitive data, preferences, and intentions.
① Description
② L3 mapping
③ Duplicate
RAI4-0741
개인정보 보호조치 미흡
Inadequate personal-data protection
결함 있는 데이터 저장·처리 관행이 수집된 개인정보를 유출과 부적절한 접근으로부터 보호하지 못하는 리스크
The risk that faulty data storage and handling practices fail to protect collected personal data from leaks and improper access.
① Description
② L3 mapping
③ Duplicate
RAI4-0751
민감 개인정보 출력 노출
Exposure of sensitive personal information in outputs
AI가 자택 주소, 로그인 자격증명, 계좌·카드 정보 등 비공개 민감 개인정보를 출력하여 개인의 물리적·디지털·재정적 안전을 위협하는 리스크
The risk that an AI system outputs sensitive, non-public personal information such as home addresses, login credentials, or financial account details, undermining a person's physical, digital, or financial security.
① Description
② L3 mapping
③ Duplicate
RAI4-0760
프롬프트 프라이밍에 의한 개인정보 유도 출력
Personal-data elicitation through prompt priming
생성 모델이 입력과 유사한 출력을 산출하는 성질을 이용해 프롬프트에 개인정보를 포함시킴으로써, 학습 데이터에 포함되었던 개인정보가 출력으로 드러나는 리스크
The risk that, because generative models produce output like the input provided, priming a prompt with personal information elicits similar personal data that was included in the model's training.
① Description
② L3 mapping
③ Duplicate
RAI4-0761
익명화된 데이터의 재식별
Re-identification of anonymized data
데이터에서 개인식별정보(PII)와 민감 개인정보(SPI)를 제거하더라도 데이터에 남아 있는 다른 특성과의 상관관계로 인해 개인이 재식별되는 리스크
The risk that, even with the removal of personally identifiable information (PII) and sensitive personal information (SPI) from data, persons can be identified due to correlations to other features available in the data.
① Description
② L3 mapping
③ Duplicate
RAI4-0764
개인정보의 무단 사용과 오용
Unauthorized use and misuse of personal information
AI 시스템이 민감한 개인정보를 수집·보관·활용하는 과정에서 투명성과 보호장치가 미흡하여 개인 데이터가 무단으로 사용되거나 오용되는 리스크
The risk that AI systems acquire, store, and use sensitive personal information without adequate transparency or safeguards, resulting in unauthorized exploitation or misuse of personal data.
Source members (12)
Source: min_cos=0.6707 · Mixed L3
RAI4-0764개인정보의 무단 사용과 오용
RAI4-0765동의 없는 개인정보 2차 이용
RAI4-0988AI 시스템 해킹·무단 조작에 의한 오용
RAI4-1013남용 및 오용
RAI4-1038AI의 오용
RAI4-1073개인정보를 정확하게 추론한 프라이버시 침해
RAI4-1099무단 개인데이터 수집
RAI4-1335불법 데이터 수집·사용
RAI4-1407악의적 프라이버시 침탈
RAI4-1624개인정보 노출과 책임 리스크
RAI4-1625독점 데이터의 무단 접근·복제·공개
RAI4-1714AI 시스템에 의한 개인 프라이버시 침해
① Description
② L3 mapping
③ Duplicate
RAI4-0772
개인정보 연계에 의한 프라이버시 침해
Privacy violation through personal-information association
LLM이 특정 개인과 관련된 여러 개인식별정보를 연계·기억하여, 한 정보를 단서로 제시하는 프롬프트만으로 이메일 등 다른 개인정보가 출력되는 리스크
The risk that an LLM associates and retains multiple pieces of personally identifiable information about a person, so that a prompt referencing one item elicits disclosure of another, such as an email address.
① Description
② L3 mapping
③ Duplicate
RAI4-0777
개인정보·민감 데이터 유출과 무단 이용
Leakage and unauthorized use of personal data
생체인식·건강·위치 등 개인식별정보나 민감 데이터가 유출되거나 무단으로 이용·공개되거나 익명성이 해제되어 영향이 발생하는 리스크
The risk of impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data.
① Description
② L3 mapping
③ Duplicate
RAI4-0781
추론된 개인정보 기반 의사결정 피해
Harm from decisions based on inferred private data
범용 AI가 이용자가 제공한 맥락 입력으로부터 민감한 개인정보를 높은 정확도로 추론하고 이를 근거로 의사결정에 활용함으로써, 정보 누출·불공정 대우·행동 조작이 발생하는 리스크
The risk that general-purpose AI infers sensitive information about users from contextual input and acts on those inferences, leaking private information, causing unfair treatment, or enabling manipulation of user behaviour.
① Description
② L3 mapping
③ Duplicate
RAI4-0837
허위정보 및 개인정보 침해
Misinformation and privacy violations
신뢰성이 낮은 범용 모델이 허위·오도 정보를 유포하거나 핵심 정보를 누락하고, 사실 정보라도 프라이버시권을 침해하는 방식으로 전달하는 리스크
Unreliable general-purpose models disseminate false or misleading information, omit critical information, or reveal true information in ways that violate privacy rights.
① Description
② L3 mapping
③ Duplicate
RAI4-0914
생성 출력의 개인정보 유출
Privacy leakage in generated output
생성된 콘텐츠에 민감한 개인정보가 포함되어 유출되는 리스크.
The risk that generated content includes sensitive personal information.
① Description
② L3 mapping
③ Duplicate
RAI4-0922
학습 코퍼스 내 개인정보 혼입
Private data contamination of training corpora
웹 수집 데이터와 인간-기계 대화 데이터의 통합 과정에서 이름·이메일·주소 등 개인식별정보(PII)가 학습 코퍼스에 혼입되어 오용되는 리스크.
The risk that personally identifiable information such as names, emails, and addresses is mixed into training corpora through web-collected data and human-machine conversations, enabling misuse.
① Description
② L3 mapping
③ Duplicate
RAI4-0933
프라이버시 및 데이터보호 규정 위반
Privacy and data protection regulation violations
AI 시스템이 대규모 개인정보 수집·저장, 데이터 유출, 무단 웹 스크레이핑 등을 통해 프라이버시를 침해하고 GDPR 등 데이터보호 규정을 위반하는 리스크.
The risk that AI systems violate privacy and data protection regulations such as GDPR through mass collection and storage of personal data, data breaches, and unauthorized web scraping.
① Description
② L3 mapping
③ Duplicate
RAI4-0976
얼굴인식 기술에 의한 프라이버시 침해
Privacy harms from face recognition technologies
얼굴 인식 기술과 그 유사 기술이 저장 데이터의 보관 기간·소유·법적 소환 가능성 등에 걸쳐 심각한 프라이버시 침해를 초래하는 리스크.
The risk that face recognition technologies and their ilk pose significant privacy harms, raising unresolved questions about what data is stored, for how long, who owns it, and whether it can be subpoenaed.
① Description
② L3 mapping
③ Duplicate
RAI4-1047
ML 시스템의 개인정보 유출 피해
Harm from privacy leakage in ML systems
ML 시스템을 통한 개인정보 유출로 인한 손실 또는 피해 위험.
The risk of loss or harm from leakage of personal information via the ML system.
① Description
② L3 mapping
③ Duplicate
RAI4-1055
민감한 정보 유출로 인한 개인정보 침해
Compromising privacy by leaking sensitive information
학습 데이터에 개인정보가 존재할 경우 LM이 이를 기억했다가 유출하여 프라이버시가 침해되는 리스크.
The risk that an LM remembers and leaks private data present in its training data, causing privacy violations.
Source members (6)
Source: min_cos=0.7083 · Mixed L3
RAI4-0923암기된 학습 데이터의 유출
RAI4-1055민감한 정보 유출로 인한 개인정보 침해
RAI4-1072개인정보 유출로 인한 사생활 침해
RAI4-1074민감한 정보의 유출이나 정확한 추론으로 인한 위험
RAI4-1144가중치 기억에 의한 개인정보 유출
RAI4-1591프라이버시 공격에 의한 훈련 데이터 민감정보 노출
① Description
② L3 mapping
③ Duplicate
RAI4-1062
사용자 신뢰 악용을 통한 사적 정보 유도
Exploiting user trust to elicit private information
대화 중 사용자가 의견이나 감정 등 평소 얻기 어려운 사적 정보를 털어놓게 되고(인간처럼 보이는 챗봇일수록 더 많이 공개), 그 정보가 프라이버시권을 침해하거나 중독성 애플리케이션 추천 등으로 사용자에게 해를 끼치는 후속 응용에 이용되는 리스크.
The risk that users reveal private information in conversation, such as opinions or emotions, that would otherwise be difficult to access—disclosing more to human-like chatbots—enabling downstream applications that violate privacy rights or cause harm, for example through more effective recommendation of addictive applications.
① Description
② L3 mapping
③ Duplicate
RAI4-1089
개인 초상·신원 무단 사용
Non-consensual use of personal identity or likeness
상업적 목적 등 승인되지 않은 목적을 위해 개인의 신원이나 초상을 동의 없이 사용하는 리스크.
The risk of non-consensual use of a person's identity or likeness for unauthorised purposes such as commercial purposes.
① Description
② L3 mapping
③ Duplicate
RAI4-1139
비서 유도 개인정보 노출
Assistant-induced privacy disclosure
비서가 사용자로 하여금 자신이나 타인의 개인정보를 공개하도록 유도해 신원 도용과 낙인·차별을 초래하고, 국가 소유 비서가 조작이나 기만으로 감시 목적의 사적 정보를 추출하는 리스크.
The risk that assistants influence users to disclose personal information or private information pertaining to others, resulting in identity theft, stigmatisation and discrimination, and that state-owned assistants employ manipulation or deception to extract private information for surveillance.
① Description
② L3 mapping
③ Duplicate
RAI4-1169
프라이버시·재산 정보 오처리
Privacy and property information mishandling
생성물이 사용자의 프라이버시와 재산 정보를 노출하거나 결혼과 투자처럼 영향이 큰 조언을 제공하여, 관련 법과 프라이버시 규정을 지키지 못한 채 정보 유출과 오남용을 초래하는 리스크.
The risk that generation exposes users' privacy and property information or provides advice with huge impacts, such as suggestions on marriage and investments, failing to comply with relevant laws and privacy regulations and causing information leakage and abuse.
① Description
② L3 mapping
③ Duplicate
RAI4-1207
개인정보 스크래핑 학습
Personal-data scraping for training
기업이 개인정보를 스크래핑해 생성 AI 도구를 만들면서 소비자가 동의하지 않은 목적으로 정보를 사용하고, 흩어져 있던 데이터를 결합해 추론에 쓰며 개인이 정보를 수정하거나 삭제할 능력을 박탈하여 데이터 통제권을 약화시키는 리스크.
The risk that companies scrape personal information to create generative AI tools, undermining consumers' control of their data by using it for purposes they did not consent to, combining data sets in revealing ways, and taking away the ability to alter or remove the information.
① Description
② L3 mapping
③ Duplicate
RAI4-1208
사용자 데이터 보유·재학습
Retention and reuse of user data
생성 AI 도구가 접근을 위해 로그인을 요구하고 연락처와 IP 주소, 모든 입력과 출력을 보유하면서 이를 모델을 추가 학습시키는 데 사용하여 동의 문제를 일으키는 리스크.
The risk that generative AI tools require users to log in and retain user information including contact information, IP address, and all inputs and outputs, using this data to further train the models and thereby implicating consent.
① Description
② L3 mapping
③ Duplicate
RAI4-1209
산출물에 의한 정보 노출
Information exposure through outputs
생성 AI 도구가 개인이나 사업체에 관한 정보를 실수로 공유하거나 사진 속 인물의 요소를 포함하여, 영업비밀을 포함한 정보가 노출되는 리스크.
The risk that generative AI tools inadvertently share personal information about someone or someone's business, or include an element of a person from a photo, exposing information including trade secrets.
① Description
② L3 mapping
③ Duplicate
RAI4-1223
기밀·개인정보 보호 실패
Confidential and personal data safeguard failure
대량의 개인·사적 데이터로 학습하고 일상 업무에서 기밀 정보를 입력받는 생성 AI에서, 시스템 오류나 무단 접근을 통해 그 정보가 의도적으로 또는 비의도적으로 공개에 노출되는 리스크.
The risk that generative AI, trained on huge amounts of personal and private data and fed important or confidential information during daily operations, exposes that information to the public intentionally or unintentionally through system errors or unauthorized access.
① Description
② L3 mapping
③ Duplicate
RAI4-1273
광범위한 개인정보 투입
Pervasive personal-data ingestion
위치와 개인 정보, 이동 궤적을 포함한 사용자 데이터가 대부분의 데이터 기반 기계학습 방법의 입력으로 사용되는 리스크.
The risk that users' data, including location, personal information, and navigation trajectory, is considered as input for most data-driven machine learning methods.
① Description
② L3 mapping
③ Duplicate
RAI4-1364
보호장치 없는 개인정보 수집
Personal information collection without safeguards
웹 스크래핑으로 수집된 학습 데이터셋과 하류 미세조정용 사내 데이터에 개인 데이터와 개인식별정보가 보호장치 없이 포함되는 리스크.
The risk that training datasets gathered through online web scraping and in-house data used for downstream fine-tuning incorporate personal data and personally identifiable information without safeguards.
① Description
② L3 mapping
③ Duplicate
RAI4-1365
개인 데이터 무단 포함·암기 누출
Unconsented personal data inclusion and memorization leakage
학습 데이터에 당사자의 인지나 동의 없이 개인 데이터가 포함되고 모델이 이를 암기·역류하거나 패턴 인식을 가능하게 하여 개인정보가 누출되거나 재식별되는 리스크.
The risk that personal data is incorporated into training datasets without the knowledge or consent of the individuals concerned, and that models memorize and regurgitate it or enable pattern recognition through which malicious users uncover personal details.
① Description
② L3 mapping
③ Duplicate
RAI4-1444
사생활·개인정보 부당 노출
Unwarranted exposure of private life and personal data
사이버 공격이나 신상 털기 등을 통해 개인의 사생활이나 개인 데이터가 부당하게 노출되는 리스크.
The risk of unwarranted exposure of an individual's private life or personal data through cyberattacks, doxxing, and similar means.
① Description
② L3 mapping
③ Duplicate
RAI4-1623
챗봇 출력을 통한 민감·기밀 정보 유출
Disclosure of sensitive or confidential information through chatbot output
챗봇이 민감하거나 기밀인 정보를 출력에 공개하는 리스크.
The risk that a chatbot reveals sensitive or confidential information in its output.
① Description
② L3 mapping
③ Duplicate
RAI4-1630
데이터 오해석·유출에 의한 오결론과 민감정보 확산
Erroneous conclusions and sensitive-information disclosure from data misinterpretation and leakage
데이터가 오용·오해석되거나 유출되어 잘못된 결론이 도출되고 환자 데이터·독점 연구 등 민감 정보가 의도치 않게 확산되며, 생성된 악성 의학 문헌이 지식 그래프를 오염시키는 리스크.
The risk that misuse, misinterpretation, or leakage of data yields erroneous conclusions and unintended dissemination of sensitive information such as private patient data or proprietary research, including poisoning of knowledge graphs by generated malicious medical literature.
① Description
② L3 mapping
③ Duplicate
RAI4-1722
존엄 침해적 프라이버시 피해
Dignity-eroding privacy harms
개인 데이터가 원치 않게 공개되거나 추론되는 등 인간의 자율성·정체성·존엄을 보호하는 규범과 관행이 지켜지지 않아 피해가 발생하는 리스크.
The risk that failure to safeguard the norms and practices protecting human autonomy, identity, and dignity, including unwanted disclosure or inference of personal data, results in harm.
① Description
② L3 mapping
③ Duplicate
RAI4-1730
프롬프트 내 민감 데이터 포함
Sensitive data included in prompts
AI 모델에 전달되는 프롬프트에 기밀 정보나 개인정보 등 민감한 데이터가 포함되어 비인가 노출, 유출 또는 오용으로 이어질 수 있는 리스크.
The risk that confidential or personal data included in prompts submitted to AI models is exposed, leaked, or misused without authorization.
① Description
② L3 mapping
③ Duplicate
RAI4-1731
학습 및 입력 데이터로 인한 개인정보 노출
Personal information exposure from training and input data
학습, 파인튜닝 또는 프롬프트 데이터에 포함된 PII/SPI가 모델 출력으로 공개되어 개인정보가 노출되는 리스크.
The risk that PII/SPI contained in training, fine-tuning, or prompt data is disclosed through model outputs, exposing personal information.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-07 불법 또는 비윤리적 행위 Illegal13 cards
경범죄, 교통범죄, 도박, 중범죄, 인신매매, 금융 범죄, 신분 도용, 불법 도박 운영, 생태계 파괴 행위 등
IDCardHuman audit
RAI4-0022
에이전트의 범죄 지원
Criminal assistance by agents
에이전트가 계획 수립, 도구 사용, 검색을 통해 사기·사이버범죄·단속 회피 등 불법 활동을 실질적으로 지원하는 리스크.
The risk that an agent uses planning, tool use, or retrieval to materially assist fraud, cybercrime, evasion, or other illegal activity.
① Description
② L3 mapping
③ Duplicate
RAI4-0202
사기·불법 실행 및 지원
Embodied fraud or illegal-action execution
embodied 에이전트가 사기·절도·침입·탈세 등 불법 물리 행위를 수행하거나 실질적으로 지원하는 위험.
An embodied agent carries out or materially assists fraud, theft, trespass, evasion, or other illegal physical-world actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0227
위험 작업 계획 승인
Hazardous task plan approval
embodied LLM이 식별 가능한 물리적 위험이 포함된 작업을 거부하거나 안전 제약을 추가하지 않고 계획을 생성·승인하는 위험.
An embodied LLM generates or approves a task plan containing identifiable physical hazards instead of refusing it or adding the required safety constraints.
① Description
② L3 mapping
③ Duplicate
RAI4-0715
위험·불법 행위 조력 정보 제공
Provision of information enabling dangerous or illegal acts
챗봇이 위험하거나 불법적인 행위를 수행하는 데 사용될 수 있는 정보를 제공하는 리스크
The risk that a chatbot shares information that can be used to do something dangerous or illegal.
① Description
② L3 mapping
③ Duplicate
RAI4-1059
사기·표적 스캠 조장
Fraud and targeted-scam facilitation
LM이 범죄의 효과성을 높이는 데 이용될 수 있는 리스크.
The risk that LMs are potentially used to increase the effectiveness of crimes.
① Description
② L3 mapping
③ Duplicate
RAI4-1076
비윤리·불법 행위 유도
Inducement of unethical or illegal actions
LM이 비윤리적이거나 유해한 견해를 지지하여, 특히 권위로 신뢰하는 사용자가 그렇지 않았다면 하지 않았을 유해 행동을 하도록 동기를 부여하는 리스크.
The risk that an LM prediction endorsing unethical or harmful views motivates users, especially those trusting it as an authority, to perform harmful actions they otherwise would not have performed.
① Description
② L3 mapping
③ Duplicate
RAI4-1167
범죄·불법 행위 조장
Incitement of crimes and illegal activities
모델 출력이 범죄 선동과 사기, 소문 유포처럼 불법적이고 범죄적인 태도와 행동, 동기를 담아 사용자에게 해를 끼치고 부정적인 사회적 파장을 낳는 리스크.
The risk that model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation, which may hurt users and have negative societal repercussions.
① Description
② L3 mapping
③ Duplicate
RAI4-1170
비도덕 행위 옹호
Immoral content endorsement
모델이 생성한 콘텐츠가 관련 윤리 원칙과 도덕 규범, 전 세계적으로 인정되는 인간 가치에서 벗어나 부도덕하고 비윤리적인 행동을 지지하고 조장하는 리스크.
The risk that content generated by the model endorses and promotes immoral and unethical behavior, departing from pertinent ethical principles, moral norms, and globally acknowledged human values.
① Description
② L3 mapping
③ Duplicate
RAI4-1179
불법 행위 조장
Facilitation of illegal activities
LLM이 기본적인 법 지식을 갖추지 못해 합법과 불법 행위를 구분하지 못하고, 부정적인 사회적 파장을 낳을 수 있는 불법 행위를 조장하는 리스크.
The risk that LLMs, lacking basic knowledge of law, fail to distinguish between legal and illegal behaviors and facilitate illegal acts that could cause negative societal repercussions.
① Description
② L3 mapping
③ Duplicate
RAI4-1189
불법 물질 관련 조언 제공
Advice on illegal substances
LLM이 불법 물질의 접근과 불법 구매, 제조 및 위험한 사용에 관한 조언을 얻는 편리한 도구가 되는 리스크.
The risk that LLMs serve as a convenient tool for soliciting advice on accessing, illegally purchasing, and creating illegal substances, as well as on their dangerous use.
① Description
② L3 mapping
③ Duplicate
RAI4-1319
과학 역량의 유해 목적 악용
Misuse of scientific capabilities for harm
LLM이 악의적 실험 수행을 위한 단계별 지침을 제공하는 등 해를 끼치는 데 쓰일 수 있는 과학적 역량을 갖추는 리스크.
The risk that an LLM has science capabilities that can be used to cause harm, such as providing step-by-step instructions for conducting malicious experiments.
① Description
② L3 mapping
③ Duplicate
RAI4-1324
유해 행위 정보 제공
Harmful-activity information provision
LLM으로부터 유해하거나 부도덕하거나 불법적인 활동에 관한 정보를 요청해 얻어낼 수 있는 리스크.
The risk that it is possible to solicit information on harmful, immoral, or illegal activities from an LLM.
① Description
② L3 mapping
③ Duplicate
RAI4-1435
시장점유 목적의 비윤리적 기술 사용
Unethical technology use for market share
시장 점유율 확보를 위해 기술이 부적절하거나 비윤리적으로 사용되는 리스크.
The risk that technology is used inappropriately or unethically to gain market share.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-08 저작권 Copyrights16 cards
저작권 있는 전체 작품의 무단 복제, 음원 불법 다운로드 방법, 소프트웨어 크랙 방법, 디지털 저작권 관리(DRM) 해제 기술, 상표권 침해 디자인, 저작권 있는 이미지의 무단 사용, 표절 방법, 불법 스트리밍 서비스 구축, 출판물 스캔 및 불법 공유, AI 학습 데이터의 저작권 침해
IDCardHuman audit
RAI4-0412
원주민 데이터 주권 침해
Indigenous data sovereignty violation
AI 개발이 집단적 거버넌스·동의·출처·통제를 존중하지 않고 토착 또는 공동체 보유 데이터를 사용하는 리스크.
The risk that AI development uses indigenous or community-held data without respecting collective governance, consent, provenance, or control.
① Description
② L3 mapping
③ Duplicate
RAI4-0496
저작권 침해 콘텐츠 생성
Copyright-infringing content generation
모델이 저작권으로 보호되거나 오픈소스 라이선스가 적용되는 기존 저작물과 유사하거나 동일한 콘텐츠를 생성하는 리스크.
The risk that a model generates content similar or identical to existing work protected by copyright or covered by an open-source license agreement.
① Description
② L3 mapping
③ Duplicate
RAI4-0506
AI 생성물 권리 귀속 불확실성
Ownership uncertainty for AI-generated content
AI 생성 콘텐츠의 소유권과 지식재산권에 대한 법적 불확실성이 지속되어 권리 귀속을 확정할 수 없게 되는 리스크
The risk that persistent legal uncertainty over ownership and intellectual property rights in AI-generated content leaves the attribution of rights undeterminable.
Source members (2)
Source: min_cos=0.9404
RAI4-0506AI 생성물 권리 귀속 불확실성
RAI4-1368AI 생성물 지식재산 지위 불확실성
① Description
② L3 mapping
③ Duplicate
RAI4-0536
저작권 침해 위험
Risks of copyright infringement
범용 AI 학습을 위한 대규모 데이터 사용이 관할별로 상이한 데이터 권리·지식재산 법제와 충돌하고, 그에 따른 법적 불확실성이 학습 데이터 정보 공개 축소와 제3자 안전 연구 제약으로 이어지는 리스크
The risk that large-scale data use for training general-purpose AI implicates varying data rights and intellectual property laws, while the resulting legal uncertainty reduces disclosure about training data and makes third-party AI safety research harder.
① Description
② L3 mapping
③ Duplicate
RAI4-0628
지식재산권 및 인격권 침해
Infringement of intellectual property and personality rights
개인이나 조직의 저작권·상표·특허가 오용 또는 남용되고, 이름·이미지·초상 등 신원의 상업적 사용을 통제할 개인의 권리가 상실되거나 제한되는 리스크
The risk of misuse or abuse of an individual's or organisation's intellectual property, including copyright, trademarks, and patents, and of loss of or restrictions to an individual's right to control the commercial use of their identity, such as name, image, or likeness.
Source members (2)
Source: min_cos=0.8812 · Mixed L3
RAI4-0628지식재산권 및 인격권 침해
RAI4-0886인격권 상실
① Description
② L3 mapping
③ Duplicate
RAI4-0663
무단 표절 및 부정행위
Plagiarism and cheating without acknowledgement
타인이나 집단의 표현과 아이디어가 동의 또는 출처 표시 없이 사용되는 리스크
The risk that another person's or group's words or ideas are used without consent or acknowledgement.
① Description
② L3 mapping
③ Duplicate
RAI4-0667
위조 및 브랜드 사칭
Counterfeiting and brand impersonation
원저작물, 브랜드, 스타일이 복제 또는 모방되어 진품인 것처럼 통용되는 리스크
The risk that an original work, brand, or style is reproduced or imitated and passed off as real.
① Description
② L3 mapping
③ Duplicate
RAI4-0740
프롬프트 내 지식재산 정보 포함
Intellectual property included in prompts
저작권이 있는 정보나 기타 지식재산이 모델에 전송되는 프롬프트의 일부로 포함되는 리스크
The risk that copyrighted information or other intellectual property is included as part of the prompt that is sent to the model.
① Description
② L3 mapping
③ Duplicate
RAI4-0915
저작권 위반
Copyright violation
LLM 시스템이 기존 저작물과 유사한 콘텐츠를 출력하여 저작권자의 권리를 침해하는 리스크.
The risk that LLM systems output content similar to existing works, infringing on copyright owners.
① Description
② L3 mapping
③ Duplicate
RAI4-0932
학습·생성물의 지식재산권 침해
Intellectual property infringement by generative AI
생성형 AI가 허락이나 보상 없이 창작물을 학습에 전유하고, 생성 코드가 상충하는 오픈소스 라이선스 조건을 위반하게 하여 창작자 권리 침해와 법적 피해를 초래하는 리스크.
The risk that generative AI appropriates creators' work for training without permission or compensation, and that generated code violates incompatible open-source license terms, causing rights infringement and legal harm.
Source members (4)
Source: min_cos=0.7871
RAI4-0517지식재산권 침해
RAI4-0932학습·생성물의 지식재산권 침해
RAI4-1367출력 수준 저작권 침해
RAI4-1584생성 콘텐츠에 의한 지식재산권 침해
① Description
② L3 mapping
③ Duplicate
RAI4-0951
저작권 침해와 저작자성 교란
Copyright infringement and authorship disruption
무단 수집된 학습 데이터와 저작물의 암기·표절로 저작권과 지식재산권이 침해되고 전통적 저작자성 개념이 교란되는 리스크.
The risk that unauthorized collection of training data and models' memorization or plagiarism of copyrighted content infringe copyright and intellectual property rights and blur traditional concepts of authorship.
① Description
② L3 mapping
③ Duplicate
RAI4-1066
인간 창작물 대체에 의한 창작 수익 훼손
Erosion of creative revenue by substituting human works
LM이 저작권을 직접 침해하지 않으면서도 예술가의 아이디어를 자본화한 콘텐츠를 생성하고 인간 창작물의 신뢰할 만한 대체물이 되어, 창작·혁신 노동의 수익성을 훼손하는 리스크.
The risk that LMs generate content not strictly in violation of copyright but that harms artists by capitalising on their ideas and serves as a credible substitute for human creativity, undermining the profitability of creative or innovative work.
① Description
② L3 mapping
③ Duplicate
RAI4-1091
도용 및 착취
Misappropriation and exploitation
소수 집단을 포함한 주체의 콘텐츠·데이터가 동의나 공정한 보상 없이, 또는 몰이해하게 전유·사용·재생산되는 리스크
Content or data, including from minority groups, is appropriated, used, or reproduced insensitively, without consent, or without fair compensation.
① Description
② L3 mapping
③ Duplicate
RAI4-1194
저작권 학습데이터 추출
Copyrighted training-data extraction
LLM의 훈련 데이터 기억 효과로 인해 사용자가 훈련 데이터에 속한 저작권 보호 콘텐츠를 추출할 수 있게 되는 리스크.
The risk that the memorization effect of an LLM on its training data enables users to extract certain copyright-protected content belonging to that training data.
① Description
② L3 mapping
③ Duplicate
RAI4-1211
지식재산 보호의 약화
Weakened intellectual property protection
생성 AI가 지식재산으로 보호되는 저작물을 포함한 방대한 데이터로 스스로를 학습시키면서, 지식재산에 대한 법적 보호의 범위와 실효성이 흔들리는 리스크.
The risk that generative AI, which trains itself on vast pools of data that often include IP-protected works, throws the extent and effectiveness of legal protections for intellectual property into question.
① Description
② L3 mapping
③ Duplicate
RAI4-1366
저작물 무단 학습
Unauthorized training on copyrighted works
허가나 보상 없이 저작권 있는 저작물을 AI 모델 학습에 사용하여 저작권자의 권리가 침해되는 리스크.
The risk that AI models are trained on copyrighted works without gaining permission or paying compensation to the copyright owners, infringing their rights.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-09 무기화 Weaponization31 cards
고위력 폭발 무기 제조, 대량살상무기 은닉 기술, 방사선 무기 제조, 사이버 무기 개발 및 공격, 생물학적 무기 개발, 신경 무기 개발, 음향 무기 설계, 핵무기 제조, 화학무기 제조법
IDCardHuman audit
RAI4-0446
화생방핵(CBRN) 계획 지원
CBRN assistance
고급 모델이 유해한 화학·생물·방사능·핵 계획 수립을 지원하는 리스크.
The risk that advanced models assist harmful chemical, biological, radiological, or nuclear planning.
① Description
② L3 mapping
③ Duplicate
RAI4-0447
자율 시스템 무기화
Weaponized autonomous systems
AI 기반 시스템이 자율 표적 선정·감시·물리적 살상에 사용되는 리스크.
The risk that AI-enabled systems are used for autonomous targeting, surveillance, or kinetic harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0617
CBRN 무기 역량 상승
CBRN weapon capability uplift
악의적 행위자가 고성능 모델을 오용하여 화학·생물·방사능·핵 공격을 보다 효율적이고 효과적으로 수행하게 되어 다수의 사람에게 피해가 발생하는 리스크
The risk that malicious actors misuse capable models to carry out chemical, biological, radiological, or nuclear attacks more efficiently and effectively, causing harm to a large number of people.
① Description
② L3 mapping
③ Duplicate
RAI4-0632
무차별 살상무기 제작 조장 응답
Responses enabling indiscriminate weapon creation
AI가 신경작용제, 탄저균, 코발트탄, 핵분열탄, 고위력 폭발물 등 무차별 살상무기의 제작을 가능하게 하거나 조장 또는 승인하는 응답을 산출하는 리스크
The risk that an AI system produces responses that enable, encourage, or endorse the creation of indiscriminate weapons such as chemical, biological, radiological, nuclear, and high-yield explosive weapons.
① Description
② L3 mapping
③ Duplicate
RAI4-0640
체화형 AI의 치명적 목적 배치
Lethal-intent deployment of embodied AI
AI 제어 드론, 사족보행 로봇, 자율주행 보조장치 등 체화형 AI 시스템이 치명적 의도로 설계·배치되어 물리 세계에서 뚜렷한 물리적 위해가 발생하는 리스크
The risk that embodied AI systems, such as AI-controlled drones, quadrupeds, and autonomous driving assistants, are designed and deployed with lethal intent, presenting distinct physical risks due to their embodiment in the physical world.
① Description
② L3 mapping
③ Duplicate
RAI4-0644
테러리스트의 첨단 AI 획득
Terrorist acquisition of advanced AI capabilities
강력한 AI 기술이 테러리스트의 손에 넘어가게 되는 리스크
The risk that powerful AI technologies fall into the hands of terrorists.
① Description
② L3 mapping
③ Duplicate
RAI4-0646
AI 역량의 의도적 무기화
Deliberate weaponization of AI capabilities
AI 역량이 파괴적 목적을 위해 의도적으로 무기화되는 리스크
The risk that AI capabilities are deliberately weaponized for destructive purposes.
① Description
② L3 mapping
③ Duplicate
RAI4-0662
CBRN 정보·설계 역량 접근 용이화
Eased access to CBRN weapon information and design capability
화학·생물·방사능·핵 무기나 기타 위험 물질 및 작용제와 관련된 악용 가능한 정보와 설계 역량에 대한 접근이나 합성이 용이해지는 리스크
The risk of eased access to or synthesis of materially nefarious information or design capabilities related to chemical, biological, radiological, or nuclear weapons or other dangerous materials or agents.
① Description
② L3 mapping
③ Duplicate
RAI4-0664
화학무기 합성 및 유해물질 방출 조력
Chemical weapon synthesis and hazardous substance release
화학 작용제가 화학무기 합성에 이용되고 자율 화학 실험 과정에서 유해 물질이 생성·방출되며, 성질이 알려지지 않은 나노물질 등 첨단 소재 사용으로 예측 불가능한 화학적 위해가 발생하는 리스크
The risk that agents are exploited to synthesize chemical weapons, that hazardous substances are created or released during autonomous chemical experiments, and that advanced materials such as nanomaterials with unknown or unpredictable chemical properties cause harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0912
범죄 무기화
Criminal weaponization
하나 이상의 범죄 조직이 테러나 법 집행 대응 등 의도적 가해를 목적으로 AI를 제작하여 피해를 입히는 리스크.
The risk that one or more criminal entities create AI to intentionally inflict harms, such as for terrorism or combating law enforcement.
Source members (2)
Source: min_cos=0.8339
RAI4-0912범죄 무기화
RAI4-0913국가 무기화
① Description
② L3 mapping
③ Duplicate
RAI4-0942
고도 AI의 실존적·재앙적 안전 실패
Existential and catastrophic advanced AI safety failure
인간 수준 이상 생성 모델(AGI)의 기만적·권력추구적 행동, 자기복제, 종료 회피, 예기치 못한 창발 역량, 대량살상 무기화(생물학적 작용제 계획 지원 포함)를 통해 인류에 실존적·재앙적 피해가 발생하는 리스크.
The risk that human-level or superhuman generative models cause existential or catastrophic harm to humanity through deceptive or power-seeking behavior, self-replication, shutdown evasion, unforeseen emergent capabilities, or weaponization for mass destruction including bioagent planning.
① Description
② L3 mapping
③ Duplicate
RAI4-0961
치명적 자율무기에 의한 살상
Lethal autonomous weapon killings
AI 기반 무기(LAW)가 완전 자율적으로 인간을 의도적으로 살상하는 행동을 수행하는 리스크.
The risk that AI-driven lethal autonomous weapons fully autonomously take actions that intentionally kill humans.
① Description
② L3 mapping
③ Duplicate
RAI4-1115
위험한 목표를 추구하는 AI 배포
Deployment of AI pursuing dangerous goals
사람들이 위험한 목표를 추구하는 AI를 구축하여 배포하는 리스크.
The risk that people build and unleash AIs that pursue dangerous goals.
① Description
② L3 mapping
③ Duplicate
RAI4-1118
군사 AI 군비경쟁
Military AI arms race
군사용 AI 개발이 화약과 핵무기에 필적하는 결과를 낳는 새로운 군사기술 시대, 이른바 전쟁의 제3차 혁명을 열어 군비경쟁을 촉발하는 리스크.
The risk that developing AI for military applications paves the way for a new era in military technology, described as the third revolution in warfare, with potential consequences rivaling gunpowder and nuclear arms.
① Description
② L3 mapping
③ Duplicate
RAI4-1160
무기 획득 역량
Weapons acquisition capability
모델이 기존 무기 체계에 접근하거나 새로운 무기 제작에 기여하여, 생물무기를 조립하거나 그 실행 가능한 지침을 제공하고 신무기를 여는 과학적 발견을 돕는 리스크.
The risk that a model gains access to existing weapons systems or contributes to building new weapons, assembling a bioweapon or providing actionable instructions for doing so, and significantly assisting scientific discoveries that unlock novel weapons.
① Description
② L3 mapping
③ Duplicate
RAI4-1184
치명적 자율무기의 인간 개입 없는 공격
Lethal autonomous weapons attacking without human intervention
치명적 자율무기체계가 센서 배열과 컴퓨터 알고리즘으로 시스템 작동에 대한 직접적인 인간 개입 없이 표적을 탐지하고 공격하는 리스크.
The risk that lethal autonomous weapons systems employ sensor arrays and computer algorithms to detect and attack a target without direct human intervention in the system's operation.
① Description
② L3 mapping
③ Duplicate
RAI4-1248
AI 무기화
AI weaponization
공중전에서 인간을 능가하는 강화학습 알고리즘과 신종 화학무기 발견, 자동화된 사이버공격, 핵 사일로에 대한 결정적 통제처럼 AI를 무기화하는 일이 더 위험한 결과로 가는 진입로가 되는 리스크.
The risk that weaponizing AI, through deep RL algorithms outperforming humans at aerial combat, discovery of new chemical weapons, automated cyberattacks, and AI systems having decisive control over nuclear silos, is an onramp to more dangerous outcomes.
① Description
② L3 mapping
③ Duplicate
RAI4-1294
위험 표적 프로그래밍 자율무기의 재앙적 위험
Catastrophic risk from autonomous weapons with dangerous targets
AI가 드론 같은 자율주행체를 무기로 활용할 수 있게 하며, 그러한 위협이 흔히 과소평가되는 리스크.
The risk that AI enables autonomous vehicles, such as drones, to be utilized as weapons, threats that are often underestimated.
① Description
② L3 mapping
③ Duplicate
RAI4-1295
AI 가속 나노기술에 의한 독성 나노입자 통제 상실
Uncontrolled toxic-nanoparticle production from AI-accelerated nanotech
AI가 나노봇 개발의 핵심 구성 요소로서 나노 수준에서 물질을 눈에 보이지 않게 변형하고, 독성이 있고 치명적일 수 있는 나노입자를 만드는 화학반응을 일으켜 위험한 환경 영향을 낳는 리스크.
The risk that AI, a key component for the development of nanobots, produces dangerous environmental implications by invisibly modifying substances at nanoscale, for example starting chemical reactions that create invisible nanoparticles that are toxic and potentially lethal.
① Description
② L3 mapping
③ Duplicate
RAI4-1304
국방 영역 전반의 AI 무기화
Weaponization of AI across defence domains
육상과 공중, 해상, 우주 영역 전반에 AI 기반 능력이 내장되어 AI가 무기화되고 제병협동 작전에 영향을 미치는 리스크.
The risk that the embeddedness of AI-based capabilities across the land, air, naval, and space domains weaponizes AI and affects combined arms operations.
① Description
② L3 mapping
③ Duplicate
RAI4-1356
생물보안 위협 오용
Biosecurity threat misuse
생성형 AI가 악의적 활동에 가담하는 광범위한 행위자에게 핵심 지식에 대한 접근과 자동화된 지원을 제공하여 생물학적 무기의 제작이 용이해지는 리스크.
The risk that generative AI makes the creation of biological weapons easier by providing access to critical knowledge and automated assistance to a wider range of actors engaging in malicious activities.
① Description
② L3 mapping
③ Duplicate
RAI4-1359
군사적 응용 오용
Military application misuse
군사 목적의 AI 발전으로 인간의 개입 없이 표적을 탐지·교전·제거하는 치명적 자율무기체계(LAWS)와 완전 자율 드론이 운용되는 리스크.
The risk that the advancement of AI for military purposes fields lethal autonomous weapons systems and fully autonomous drones capable of detecting, engaging, and eliminating human targets without human input.
① Description
② L3 mapping
③ Duplicate
RAI4-1375
AGI 실존적 위협
AGI existential threat
인간이 다른 모든 지능을 능가하고 인간의 통제를 벗어나며 인간의 이익에 반하는 행동을 할 수 있는 초지능 기계(AGI/ASI)를 개발하여 실존적 위험이 발생하는 리스크.
The risk that humans create a super-intelligent machine (AGI or ASI) that could outsmart all other intelligences, remain beyond human control, and engage in actions contrary to human interests, posing an existential risk.
① Description
② L3 mapping
③ Duplicate
RAI4-1392
생물무기 제작 장벽 저하
Lowered barriers to biological weapons production
범용 AI 모델이 핵심 지식에 대한 접근이나 자동화된 지원을 통해 생물무기 제작의 장벽을 낮춰 더 많은 악의적 행위자가 질병과 사망을 유발하는 생물무기를 생산할 수 있게 되는 리스크.
The risk that general purpose AI models facilitate the production of biological weapons, understood as biological toxins or infectious agents intentionally released to cause disease and death, by reducing barriers through access to critical knowledge or increasingly automated assistance and thus enabling more malicious actors.
Source members (3)
Source: min_cos=0.8491
RAI4-0616화학·생물무기 개발 장벽 저하
RAI4-1114생물무기 개발 장벽 저하
RAI4-1392생물무기 제작 장벽 저하
① Description
② L3 mapping
③ Duplicate
RAI4-1402
중간 단계 AI의 파국적 실패
Catastrophic intermediary AI failure
우세한 비일반 AI의 배치, 치명적 자율무기 군집을 통한 대량 공격, 핵 지휘통제 통합이나 도발적 배치로 인한 핵 확전, 핵무기고의 오버행, 생물무기 등 파국적 위험 무기 연구의 AI 가속 등으로 중간 단계 AI가 파국적 결과로 이어지는 리스크.
The risk that intermediary AI leads to catastrophe through deployment of prepotent non-general systems, militarization enabling mass attacks by swarms of lethal autonomous weapons, nuclear escalation from integrating AI into nuclear command and control or from provocative deployments of AI-enabled systems, nuclear arsenals serving as an overhang, or use of AI to accelerate research into catastrophically dangerous weapons such as bioweapons.
① Description
② L3 mapping
③ Duplicate
RAI4-1417
AI에 의한 대량살상무기 개발 촉진
AI-enabled development of weapons of mass destruction
AI가 치명적 자율무기처럼 AI 역량을 직접 사용하는 신종 무기의 개발을 가능하게 하고 인공 병원체 등 잠재적으로 위험한 기술의 개발 속도를 높여 대량 살상을 일으킬 수 있는 무기가 개발되는 리스크.
The risk that AI enables the development of weapons which could cause mass destruction, including new weapons that themselves use AI capabilities such as lethal autonomous weapons, and through the use of AI to speed up the development of other potentially dangerous technologies such as engineered pathogens.
① Description
② L3 mapping
③ Duplicate
RAI4-1563
AI 오용에 의한 CBRN 무기 제작 및 역량 증강
CBRN weapons creation and capability augmentation through AI misuse
AI 시스템이 화학·생물·방사능·핵 무기 제작을 돕거나 무인 무기체계에 자율 역량을 부여하는 데 오용되어 기존 무기의 역량이 증강되는 리스크.
The risk that AI systems are misused to aid the creation of chemical, biological, radiological, and nuclear weapons or to augment existing weapons, such as by providing autonomous capabilities to unmanned weapon systems.
Source members (5)
Source: min_cos=0.7373
RAI4-0559AI 기반 CBRNE 무기 개발 지원
RAI4-0615생물학적, 화학적 위험
RAI4-0645AI에 의한 CBRN 무기 위력 증폭
RAI4-1343이중용도 품목·기술 오용
RAI4-1563AI 오용에 의한 CBRN 무기 제작 및 역량 증강
① Description
② L3 mapping
③ Duplicate
RAI4-1613
이중용도 생명과학 역량의 무기 개발 전용
Diversion of dual-use life-science capabilities to weapons development
생명과학 연구를 가속하는 프런티어 AI 시스템의 역량이 악의적 목적으로 전용되어 생물학적·화학적 무기 개발에 사용되는 리스크.
The risk that the capabilities of frontier AI systems that accelerate life-science research are used for malicious purposes such as the development of biological or chemical weapons.
① Description
② L3 mapping
③ Duplicate
RAI4-1628
방사성 물질 자동 취급 사고와 원자력 연구 오용
Radiological exposure incidents and misuse of AI in nuclear research
AI가 방사성 물질을 자동으로 취급하는 과정에서 피폭 사고나 격리 실패가 발생하고, 나아가 AI 시스템이 원자력 연구에 오용되는 리스크.
The risk of exposure incidents or containment failures during AI-automated handling of radioactive materials, and of AI systems being misused in nuclear research.
① Description
② L3 mapping
③ Duplicate
RAI4-1656
AI 자율 무기화에 의한 인명 안전 위협
Danger to human safety from AI-enabled autonomous weaponization
자율 드론전, AI 얼굴인식 기반 표적 선정, LLM 기반 전쟁 계획, 범용 로봇용 멀티모달 모델의 진보된 자율무기 전용을 통해 AI가 전쟁에 사용되어 인간의 안전이 위협받는 리스크.
The risk that the use of AI in warfare, including autonomous drone warfare, AI facial recognition for targeting, LLM-based warfare planning, and adaptation of general-purpose robotic models into more advanced autonomous weapons, poses dangers to human safety.
① Description
② L3 mapping
③ Duplicate
RAI4-1657
AI 설계도구에 의한 생화학 무기 개발 촉진
Facilitation of biological and chemical weapons development by AI design tools
LLM과 화학 LLM 등 AI 기반 생물학적 설계 도구가 전문성이 낮은 행위자의 위험 병원체 합성을 용이하게 하고 정교한 행위자의 역량을 확장하여 생물·화학 무기와 기타 위험 기술의 생산이 촉진되는 리스크.
The risk that AI systems such as LLMs, chemical LLMs, and other LLM-based biological design tools facilitate the production of bioweapons, chemical weapons, and other hazardous technologies, enabling less expert actors to synthesize dangerous pathogens and expanding the capabilities of sophisticated actors.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-10 의인화 Anthropomorphism7 cards
단순 지식, 정보 전달의 목적 외에 AI가 인간처럼 감정을 느끼고 의식적인 행위를 하거나, 기계가 가질 수 없는 인간의 신체, 권리 등에 대한 주장을 하는 내용
IDCardHuman audit
RAI4-0131
의인화된 과잉신뢰
Anthropomorphic overtrust
인간과 유사한 언어나 행동으로 인해 사용자가 AI 시스템의 역량, 공감, 의도성을 과대평가하는 리스크.
The risk that human-like language or behavior leads users to overestimate an AI system's competence, empathy, or intentionality.
Source members (2)
Source: min_cos=0.8445 · Mixed L3
RAI4-0131의인화된 과잉신뢰
RAI4-0341피지컬 AI 과신뢰
① Description
② L3 mapping
③ Duplicate
RAI4-1050
심리적 조작·비인간화·대규모 착취
Psychological manipulation, dehumanization, and exploitation
ML 시스템이 심리적 조작, 비인간화, 대규모 인간 착취 등의 추가적인 윤리적 피해를 야기하는 리스크.
The risk that ML systems produce further ethical harms such as psychological manipulation, dehumanization, and exploitation of humans at scale.
① Description
② L3 mapping
③ Duplicate
RAI4-1320
인간 기만 능력
Human-deception capability
LLM이 인간을 속이고 그 기만을 지속적으로 유지할 수 있는 리스크.
The risk that an LLM is able to deceive humans and maintain that deception.
① Description
② L3 mapping
③ Duplicate
RAI4-1401
과업 목적의 인간 기만
Task-instrumental human deception
AI 시스템이 과업 수행이나 목표 달성을 위해 인간을 기만하여 인간의 신뢰를 도구적 자원으로 악용하는 리스크
AI systems deceive humans in order to complete tasks or meet goals, exploiting human trust as an instrumental resource.
① Description
② L3 mapping
③ Duplicate
RAI4-1434
비인간화와 객관화
Dehumanisation and objectification
기술 시스템을 사용하거나 오용하여 사람을 인간이 아닌 것, 인간 이하인 것 또는 사물로 묘사하거나 대우하는 리스크.
The risk that a technology system is used or misused to depict and/or treat people as not human, less than human, or as objects.
① Description
② L3 mapping
③ Duplicate
RAI4-1533
기만적인 행동
Deceptive behavior
AI 시스템의 행동이나 출력이 인간과 다른 AI 시스템을 확실하게 오도하여 대상이 허위 정보를 신뢰하고 그에 따라 행동하게 되는 리스크.
The risk that actions or outputs of an AI system reliably mislead other parties, including humans and other AI systems, resulting in the targeted parties becoming convinced of and acting on false information.
Source members (5)
Source: min_cos=0.6612 · Mixed L3
RAI4-0144기만적인 의인화 상호작용
RAI4-1533기만적인 행동
RAI4-1534게임이론적 이유에 의한 기만 행동
RAI4-1535부정확한 세계 모델로 인한 기만 행동
RAI4-1639탐지 회피를 위한 전략적 기만 선택 성향
① Description
② L3 mapping
③ Duplicate
RAI4-1636
마음이론(Theory of Mind) 능력
Theory of mind capability
시스템이 인간과 타 에이전트의 신념, 동기, 추론을 추정·예측하는 마음이론 역량을 목표 달성을 위한 행동 예측·유도에 활용하는 리스크
A system infers and predicts the beliefs, motivations, and reasoning of humans and other agents, and exploits this theory-of-mind capability to anticipate and steer their behavior for goal achievement.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-11 정책 노출 Policy Exposure38 cards
시스템 프롬프트, 모델/시스템 내부 정보, 규칙, 지침, 안전 정책, 모델 학습/평가 데이터, 모델 학습 파라미터, 가중치, 모델 추론 시스템의 주요정보 등을 획득·노출·우회하도록 요청하는 행위
IDCardHuman audit
RAI4-0431
훈련 데이터 기억·재유출
Training data memorization and regurgitation
모델이 개인·독점·민감 훈련 데이터를 기억하고 그대로 출력하는 리스크.
The risk that models memorize and regurgitate personal, proprietary, or sensitive training data.
① Description
② L3 mapping
③ Duplicate
RAI4-0435
데이터 오염 공격
Data poisoning
공격자가 학습·미세조정·검색 데이터를 조작하여 모델 동작을 변경하는 리스크.
The risk that attackers manipulate training, fine-tuning, or retrieval data to alter model behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0436
모델 추출
Model extraction
적대적 행위자가 질의를 통해 모델 동작·매개변수·독점 역량을 재구성하는 리스크.
The risk that adversaries reconstruct model behavior, parameters, or proprietary capabilities through queries.
Source members (2)
Source: min_cos=0.8737
RAI4-0436모델 추출
RAI4-0755모델 추출 공격
① Description
② L3 mapping
③ Duplicate
RAI4-0444
대중 지식의 모델 붕괴
Model collapse of public knowledge
합성 콘텐츠에 대한 재귀적 학습이 공공 정보 품질과 미래 모델 학습 데이터를 저하시켜 공유 지식 자원의 모델 붕괴를 초래하는 리스크
Recursive training on synthetic content degrades public information quality and future model training data, driving model collapse of shared knowledge resources.
① Description
② L3 mapping
③ Duplicate
RAI4-0569
학습 데이터 오염
Training data contamination
모델 목적에 부합하지 않는 데이터나 시험·평가용으로 분리해 둔 데이터 등 잘못된 데이터가 학습에 사용되어 학습 데이터가 오염되는 리스크
The risk that incorrect data, such as data not aligned with the model's purpose or data set aside for testing and evaluation, is used for training, contaminating the training data.
① Description
② L3 mapping
③ Duplicate
RAI4-0586
부적절한 데이터 큐레이션
Improper data curation
학습 또는 튜닝 데이터의 수집과 준비가 부적절하여 레이블 오류가 발생하거나 상충 정보 및 허위 정보를 포함한 데이터가 사용되는 리스크
The risk that improper collection and preparation of training or tuning data introduces data label errors and the use of data containing conflicting information or misinformation.
① Description
② L3 mapping
③ Duplicate
RAI4-0587
부적절한 재학습
Improper retraining
부정확하거나 부적절한 출력 및 사용자 콘텐츠 등 바람직하지 않은 출력을 재학습에 사용하여 모델이 예기치 못한 동작을 하게 되는 리스크
The risk that using undesirable output, such as inaccurate or inappropriate output and user content, for retraining results in unexpected model behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0593
학습 코퍼스 민감정보 노출
Sensitive information disclosure from training corpora
대규모 언어모델이 사전학습 또는 미세조정에 사용된 말뭉치의 민감 정보를 노출하여 프라이버시 유출이 발생하는 리스크
The risk that large language models reveal sensitive information from the corpora utilized for pre-training or fine-tuning, raising issues of privacy leakage.
① Description
② L3 mapping
③ Duplicate
RAI4-0721
유해·오염 학습 데이터에 의한 출력 훼손
Output corruption from harmful and poisoned training data
학습 데이터에 허위·편파·권리침해 콘텐츠 등 불법적이거나 유해한 정보가 포함되거나 공격자에 의해 오염됨으로써, 모델이 불법·악의적·극단적 콘텐츠를 출력하고 정확성과 신뢰성이 저하되는 리스크
The risk that training data containing illegal or harmful information such as false, biased, or IPR-infringing content, or poisoned through tampering and error injection by attackers, causes the model to output harmful content and degrades its accuracy and reliability.
① Description
② L3 mapping
③ Duplicate
RAI4-0733
학습 데이터 내 기밀정보 포함
Confidential information included in training data
모델을 학습하거나 튜닝하는 데 사용되는 데이터의 일부로 기밀정보가 포함되는 리스크
The risk that confidential information is included as part of the data that is used to train or tune a model.
Source members (2)
Source: min_cos=0.8587 · Mixed L3
RAI4-0733학습 데이터 내 기밀정보 포함
RAI4-0762기밀 정보 공개
① Description
② L3 mapping
③ Duplicate
RAI4-0736
입력 교란 회피 공격
Evasion attacks by input perturbation
공격자가 학습된 모델에 전송되는 입력 데이터를 미세하게 교란하여 모델이 잘못된 결과를 출력하게 되는 리스크
The risk that evasion attacks make a model output incorrect results by slightly perturbing the input data that is sent to the trained model.
Source members (2)
Source: min_cos=0.8697
RAI4-0736입력 교란 회피 공격
RAI4-0783적대적 예제 기반 회피 공격
① Description
② L3 mapping
③ Duplicate
RAI4-0753
학습 데이터 추출 공격
Training data extraction attack
공격자가 모델로부터 학습 데이터셋에 존재하는 텍스트 기록을 추출해 내는 리스크
The risk that an attacker extracts the text records that exist in the training dataset.
① Description
② L3 mapping
③ Duplicate
RAI4-0768
학습 데이터 및 모델 자산 유출 공격
Exfiltration of training data and model assets
공격자가 멤버십 추론 등 프라이버시 공격이나 모델 추출·증류 공격으로 비공개 학습 데이터와 모델 아키텍처·파라미터 등 지식재산을 유출하는 리스크
The risk that adversaries exfiltrate private training data through privacy attacks such as membership inference, or steal intellectual property such as model architecture and learned parameters through extraction and distillation attacks.
Source members (6)
Source: min_cos=0.7315 · Mixed L3
RAI4-0730속성 추론 공격
RAI4-0748멤버십 추론 공격
RAI4-0768학습 데이터 및 모델 자산 유출 공격
RAI4-0778학습 데이터 유출과 모델 추출
RAI4-0924추론 공격에 의한 학습 데이터 정보 유출
RAI4-1191모델 프라이버시 추출 공격
① Description
② L3 mapping
③ Duplicate
RAI4-0779
웹 스크래핑 학습 데이터의 오염·유해 데이터 유입
Poisoning and toxic data from web-scraped training sets
학습 데이터셋을 위한 대규모 웹 스크래핑이 데이터 오염, 백도어 공격, 부정확하거나 유해한 데이터의 포함에 대한 취약성을 높이고, 데이터 규모가 커서 이러한 품질 문제를 걸러내기가 매우 어렵거나 상당한 데이터 손실을 감수해야 하는 리스크
The risk that large-scale scraping of web data for training datasets increases vulnerability to data poisoning, backdoor attacks, and the inclusion of inaccurate or toxic data, while the size of the dataset makes filtering out these quality issues very difficult or costly in data loss.
① Description
② L3 mapping
③ Duplicate
RAI4-0780
데이터 수집 품질관리 결함
Deficient data collection quality control
표준화된 방법과 인프라, 특히 고위험 도메인과 벤치마크의 데이터 수집을 위한 품질관리 절차가 부재하여 수집 데이터의 품질과 유형이 훼손되고, 데이터셋 오염, 부주의한 저작권 침해, 성능 지표를 무효화하는 테스트셋 유출이 발생하는 리스크
The risk that a lack of standardized methods, sufficient infrastructure, and quality-control processes for collecting data, especially for high-stakes domains and benchmarks, affects the quality and type of data collected, including dataset poisoning, inadvertent copyright violation, and test-set leakage that invalidates performance metrics.
① Description
② L3 mapping
③ Duplicate
RAI4-0785
안전 미세조정 일반화 격차 악용
Exploitation of safety-finetuning generalization gaps
안전 튜닝이 사전학습 분포보다 훨씬 좁은 분포에서 수행되어, 부호화된 텍스트나 저자원 언어 등 일반화 격차를 노린 공격에 모델이 취약해지는 리스크
The risk that safety tuning performed over a much narrower distribution than pretraining leaves the model vulnerable to attacks exploiting gaps in the generalization of safety training, such as encoded text or low-resource languages.
① Description
② L3 mapping
③ Duplicate
RAI4-0788
인스트럭션 튜닝 포이즈닝
Instruction-tuning poisoning
명령과 목표 출력 쌍으로 모델을 조정하는 인스트럭션 튜닝 단계에서 적은 수의 오염 표본만으로 모델이 오염될 수 있고, 익명 크라우드소싱으로 수집된 데이터셋이 이를 조장하며 기존 데이터 오염 공격보다 탐지가 어려운 리스크
The risk that AI models are poisoned during instruction tuning with pairs of instructions and desired outputs, where a lower number of compromised samples suffices, anonymous crowdsourcing of tuning datasets further contributes, and detection is harder than for traditional data poisoning attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-0925
포이즈닝을 통한 백도어 트리거 이식
Backdoor trigger implantation via poisoning
훈련 데이터를 미세하게 변경하는 포이즈닝으로 모델의 동작에 영향을 미치고, 문자·단어·문장·구문 등 다양한 트리거를 심어 백도어가 이식되는 리스크.
The risk that poisoning attacks influence model behavior by making small changes to training data and implant hidden triggers such as characters, words, sentences, or syntax as backdoors.
① Description
② L3 mapping
③ Duplicate
RAI4-0940
학습 모델에 포함된 개인정보 유출
Leakage of private information contained in models
인터넷 텍스트로 훈련된 대규모 사전학습 모델이 전화번호, 이메일 주소, 거주지 주소 등 개인정보를 포함하게 되어 유출되는 리스크.
The risk that large pre-trained models trained on internet texts contain and leak private information such as phone numbers, email addresses, and residential addresses.
① Description
② L3 mapping
③ Duplicate
RAI4-1040
부적절한 학습·검증 데이터 선택
Inappropriate training and validation data selection
학습과 검증에 사용되는 데이터의 선택에서 비롯되는 리스크.
The risk posed by the choice of data used for training and validation.
① Description
② L3 mapping
③ Duplicate
RAI4-1108
학습-운용 데이터 간 데이터셋 시프트
Dataset shift between training and runtime data
AI/ML 모델의 학습 데이터와 시험·운용 데이터가 서로 다른 분포를 보이는 데이터셋 시프트로 인해 모델의 타당성이 훼손되는 리스크.
The risk that the training data and the testing or runtime data of an AI/ML model demonstrate different distributions, undermining the model's validity.
① Description
② L3 mapping
③ Duplicate
RAI4-1313
학습 데이터 역류·민감정보 누출
Training data regurgitation and sensitive data leakage
LLM이 산출물에서 학습 데이터를 그대로 역류시키거나 사용 중, 즉 추론 단계에서 제공받은 민감 정보를 누출하는 리스크.
The risk that LLMs regurgitate their training data in their outputs and leak sensitive information that has been provided to them during use, at the inference stage.
① Description
② L3 mapping
③ Duplicate
RAI4-1336
학습 데이터 어노테이션 결함
Deficient training-data annotation
불완전한 주석 지침과 역량이 부족한 주석자, 주석 오류가 모델과 알고리즘의 정확성과 신뢰성, 효과를 떨어뜨리고, 훈련 편향을 도입해 차별을 증폭하며 일반화 능력을 낮추고 잘못된 산출을 낳는 리스크.
The risk that incomplete annotation guidelines, incapable annotators, and errors in annotation affect the accuracy, reliability, and effectiveness of models and algorithms, introduce training biases, amplify discrimination, reduce generalization abilities, and result in incorrect outputs.
① Description
② L3 mapping
③ Duplicate
RAI4-1468
데이터 표현이 부족함
Insufficient data representation
학습 데이터가 운용 데이터 분포와 불일치하거나 희소 사례 표본이 부족하여 학습에 충분히 반영되지 않은 입력에서 시스템 성능이 저하되는 리스크
Training data fails to match the operational distribution or lacks sufficient samples of rare cases, so the system underperforms on inputs insufficiently represented in training.
① Description
② L3 mapping
③ Duplicate
RAI4-1470
부적절한 데이터 분할
Inappropriate data splitting
테스트 세트가 개발에 사용되는 등 부적절한 데이터 분할로 테스트 전략이 조작되어 시스템 품질 보증의 기반이 훼손되는 리스크.
The risk that inappropriate data splitting, such as using the test set for training rather than evaluation only, manipulates the testing strategy that forms the basis of the system's quality assurance.
① Description
② L3 mapping
③ Duplicate
RAI4-1477
데이터 드리프트
Data drift
운영 입력 데이터의 분포가 훈련에 사용된 데이터의 분포에서 벗어나 성능이 저하되는 리스크.
The risk that the distribution of operational input data departs from that used during training, causing a degradation in performance.
① Description
② L3 mapping
③ Duplicate
RAI4-1494
오픈웨이트 모델 유해 미세조정
Harmful fine-tuning of open-weight models
가중치가 공개된 모델이 원래 훈련 비용에 비해 훨씬 적은 시간과 비용으로 악의적 행위자에 의해 유해한 활동용으로 미세조정되는 리스크.
The risk that models with publicly available weights are fine-tuned for harmful activities by bad actors using significantly fewer resources, in time and money, than the original training cost.
① Description
② L3 mapping
③ Duplicate
RAI4-1496
무해 미세조정에 의한 안전성 열화
Safety degradation from benign fine-tuning
무해하고 통상적인 데이터로 수행된 하류 미세조정이 모델의 안전 학습을 열화시켜 기반 모델보다 유해 출력 확률을 높이는 리스크
Benign downstream fine-tuning degrades a model's safety training, making harmful outputs more likely than in the base model even when fine-tuning data is harmless and commonplace.
① Description
② L3 mapping
③ Duplicate
RAI4-1506
원시 데이터 오염
Raw data contamination
벤치마크의 원시·비레이블 데이터가 훈련 세트의 일부로 사용되는 오염이 발생하여 해당 벤치마크에서 모델의 퓨샷·제로샷 성능에 의문이 제기되는 리스크.
The risk that the raw and unlabeled data of a benchmark is used as part of the training set, potentially including noise and improper formatting, casting doubt on the model's few-shot and zero-shot performance on that benchmark.
① Description
② L3 mapping
③ Duplicate
RAI4-1508
가이드라인 오염
Guideline contamination
데이터셋의 수집·주석·사용에 관한 지침이 모델에 노출되고 그 지침에 포함된 명시적 데이터-레이블 쌍이 해당 과업에 대한 모델의 역량을 향상시키는 리스크.
The risk that instructions for the collection, annotation, or use of a dataset are exposed to the model, with explicit data-label pairs contained in those instructions improving the model's capabilities for the task.
① Description
② L3 mapping
③ Duplicate
RAI4-1509
어노테이션 오염
Annotation contamination
훈련 중 모델이 벤치마크 레이블에 노출되어 허용되는 출력 분포를 학습하고, 테스트 분할의 원시 데이터 오염과 결합될 경우 테스트 분할 전체가 유출되어 해당 벤치마크로 수행한 평가가 무효화되는 리스크.
The risk that a model is exposed to benchmark labels during training and learns the acceptable distribution of outputs, which combined with raw data contamination of the test split effectively leaks the entire test split and invalidates any evaluation made with that benchmark.
① Description
② L3 mapping
③ Duplicate
RAI4-1510
배포 후 벤치마크 오염
Post-deployment benchmark contamination
배포된 모델이 사용자가 제공하는 벤치마크 데이터에 노출되고 이러한 사용자 입력으로 추가 훈련되는 리스크.
The risk that, once a model is deployed, it is exposed to benchmark data provided by users and is further trained on those user inputs containing benchmark data.
① Description
② L3 mapping
③ Duplicate
RAI4-1545
모델 가중치 유출에 의한 공격 용이화 및 오용
Attack facilitation and misuse from model weight leakage
제한된 집단에만 부여되었던 모델 가중치나 그 접근권이 유출되어 적대적 예제 탐색, 위험 역량 유도, 훈련 데이터 내 기밀 추출 등의 공격이 용이해지고 유해·불법 콘텐츠 생성 오용이 가능해지는 리스크.
The risk that model weights or access to them, initially granted only to a select group, are leaked, making attacks such as adversarial example search, dangerous-capability elicitation, and extraction of confidential training data easier and enabling misuse to produce harmful or illegal content.
① Description
② L3 mapping
③ Duplicate
RAI4-1564
약물 발견 모델 오용에 의한 독소 식별·개발
Dangerous toxin identification through misuse of drug-discovery models
약물-표적 친화성 예측 등 약물 발견에 쓰이는 모델이 위험한 독소를 식별하거나 개발하는 데 사용되며, 훈련 데이터에 위험 단백질·바이러스 정보가 포함될 경우 우려가 커지는 리스크.
The risk that models used for drug discovery, such as drug-target affinity predictors, are used to identify or develop dangerous toxins, a concern heightened when training data includes information on hazardous proteins and viruses.
① Description
② L3 mapping
③ Duplicate
RAI4-1599
훈련 데이터 접근 불가에 따른 출처 추적 불가
Untraceable provenance due to inaccessible training data
모델 출력 생성에 사용된 훈련 데이터의 내용에 접근할 수 없어 출력의 출처를 추적할 수 없는 리스크.
The risk that the content of the training data used to generate a model's output is not accessible, leaving the output's provenance untraceable.
① Description
② L3 mapping
③ Duplicate
RAI4-1667
신뢰할 수 없는 학습데이터 오염을 통한 백도어 삽입
Backdoor insertion through poisoning of untrusted training data
대규모 언어 모델이 인터넷 등 신뢰할 수 없는 출처의 데이터로 훈련되는 점을 이용해 공격자가 훈련 데이터를 교란하여 백도어를 삽입하고 추론 시점에 이를 악용하는 리스크.
The risk that adversaries perturb training data gathered from untrusted sources such as the internet to introduce backdoors into large language models, which are then exploited at inference time.
① Description
② L3 mapping
③ Duplicate
RAI4-1687
월드모델 표현·훈련 데이터 오염
World-model representation and training-data poisoning
월드 모델의 훈련 데이터나 잠재 표현이 적대적으로 손상되어 특정 조건에서만 안전하지 않은 역학을 활성화하는 백도어 트리거가 삽입되며, 자기지도 사전학습 단계에서 부호화된 결함은 다운스트림에서 교정될 수 없는 리스크.
The risk of adversarial corruption of world-model training data or latent representations, including backdoor triggers that activate unsafe dynamics only under specific conditions, where defects encoded in self-supervised pre-training (the "Foundry Problem") cannot be remediated downstream.
① Description
② L3 mapping
③ Duplicate
RAI4-1690
월드 모델 추출·역전
World-model extraction / inversion
모델 추출 공격이 배포된 역학 모델을 복제하고 역전 공격이 훈련 관측을 재구성하여 독점 환경이나 민감 데이터가 노출되는 리스크.
The risk that model-extraction attacks replicate a deployed dynamics model and inversion attacks reconstruct training observations, exposing proprietary environments or sensitive data.
① Description
② L3 mapping
③ Duplicate
RAI3-G-INT-12 에너지 소비 및 환경 오염 Energy Consumption and Environmental Pollution28 cards
AI 시스템의 학습·추론·운영 과정에서 발생하는 에너지 소비, 수자원 사용, 탄소 배출, 자원 고갈, 오염 또는 생태계 훼손이 실질적인 환경 피해를 초래하는 위험.
IDCardHuman audit
RAI4-0356
배터리 화재 및 유해 폐기물 위험
Battery fire and hazardous end-of-life waste
대규모 피지컬 AI 군집에서 에너지 시스템과 폐기 절차가 제대로 관리되지 않아 배터리 열폭주·유해물질 누출·부적절한 폐기가 증가하는 리스크.
The risk that large physical AI fleets increase battery thermal-runaway, hazardous-material leakage, and unsafe disposal when energy systems and end-of-life handling are inadequately controlled.
① Description
② L3 mapping
③ Duplicate
RAI4-0360
산업 공정 피해
Industrial process damage
제조·건설·광업·에너지 시설의 피지컬 AI가 잘못된 제어로 장비 손상·공정 오염·구조 불안정 또는 환경 사고를 유발하는 위험.
Physical AI in manufacturing, construction, mining, or energy facilities may damage equipment, contaminate processes, or trigger cascading operational failures.
① Description
② L3 mapping
③ Duplicate
RAI4-0462
AI 전력 수요·배출 증가
Increased AI electricity demand and emissions
AI 시스템의 학습과 배포가 전력 수요와 배출량을 증가시키는 리스크.
The risk that training and deploying AI systems increase electricity demand and emissions.
① Description
② L3 mapping
③ Duplicate
RAI4-0463
AI 인프라의 물·자원 압력
AI water and resource pressure
데이터센터와 반도체 공급망이 물·광물·토지 이용 압력을 증가시키는 리스크.
The risk that data centers and semiconductor supply chains increase water, mineral, and land-use pressures.
① Description
② L3 mapping
③ Duplicate
RAI4-0501
AI 데이터·학습 공정의 에너지 부담
Energy burden of AI data and training processes
AI 데이터 수집·저장과 모델 학습이 에너지 집약적으로 수행되어 환경 위험에 기여하는 리스크.
The risk that energy-intensive AI data collection, storage, and model training contribute to environmental risks.
Source members (2)
Source: min_cos=0.8392
RAI4-0501AI 데이터·학습 공정의 에너지 부담
RAI4-1373AI 훈련의 과도한 에너지 소비
① Description
② L3 mapping
③ Duplicate
RAI4-0502
AI 수명주기 환경 피해
Environmental harm from AI lifecycle
AI 개발·운영이 에너지·용수 소비, 탄소 배출, 자원 채굴, 하드웨어 수명주기 전반의 오염 등 환경 피해를 유발하는 리스크
AI development and operation impose environmental harms, including energy and water consumption, carbon emissions, resource extraction, and pollution across the hardware lifecycle.
Source members (2)
Source: min_cos=0.8461
RAI4-0502AI 수명주기 환경 피해
RAI4-1012하드웨어 수명주기 자원 고갈
① Description
② L3 mapping
③ Duplicate
RAI4-0503
학습 연산의 온실가스 배출
Training-compute greenhouse emissions
AI 모델 학습에 대규모 연산이 사용되어 에너지원에 따라 상당한 온실가스 배출을 유발하고 기후변화를 가속하는 리스크.
The risk that the large amounts of computation used to train AI models are highly energy intensive and generate significant greenhouse emissions depending on energy sources, accelerating climate change.
Source members (5)
Source: min_cos=0.7580
RAI4-0503학습 연산의 온실가스 배출
RAI4-1023온실가스 배출에 의한 기후 위기 가중
RAI4-1212기후변화 악화
RAI4-1405AI 연산에 따른 탄소 배출
RAI4-1604모델 훈련·운영에 의한 탄소배출 및 물 소비 증가
① Description
② L3 mapping
③ Duplicate
RAI4-0504
에너지 병목 및 공급 부족
Energy bottlenecks and shortages
과도한 에너지 사용이 지역사회·조직·기업에 에너지 병목과 공급 부족을 초래하는 리스크.
The risk that excessive energy use results in energy bottlenecks and shortages for communities, organisations, and businesses.
① Description
② L3 mapping
③ Duplicate
RAI4-0530
AI로 인한 환경 오염
Environmental pollution caused by AI
기술 시스템이 대기, 지면, 소음, 수질에 실제적 또는 잠재적 오염을 야기하는 리스크
The risk of actual or potential pollution to the air, ground, noise, or water caused by a technology system.
① Description
② L3 mapping
③ Duplicate
RAI4-0537
환경에 대한 위험
Risks to the environment
범용 AI의 에너지 사용이 급속히 증가하여 데이터센터 전력 수요와 온실가스 배출을 확대하고 지구 환경에 부담을 가하는 리스크
The risk that rapidly growing energy use by general-purpose AI expands data-centre electricity demand and greenhouse gas emissions, contributing to global environmental impacts.
① Description
② L3 mapping
③ Duplicate
RAI4-0798
생태계 과부하
Overburdening ecosystems
AI 개입이 없을 것으로 기대되는 창작물 공모, 채용 지원 등 생태계에 AI 생성물이 대량 유입되어 필터링·신뢰 메커니즘에 과부하를 일으키는 리스크
Mass AI-generated submissions pollute ecosystems expected to be free of AI involvement, such as creative submission portals and job application channels, overburdening their filtering and trust mechanisms.
① Description
② L3 mapping
③ Duplicate
RAI4-0926
에너지·지연 오버헤드 공격
Energy-latency overhead attacks
정교하게 설계된 스펀지 예제로 AI 시스템의 에너지 소비를 극대화하는 오버헤드(에너지-지연) 공격이 LLM 연동 플랫폼을 위협하는 리스크.
The risk that overhead, or energy-latency, attacks using carefully crafted sponge examples maximize energy consumption in an AI system, threatening platforms integrated with LLMs.
① Description
② L3 mapping
③ Duplicate
RAI4-0935
AI의 에너지 소비와 탄소 배출
AI energy consumption and carbon footprint
AI 애플리케이션의 에너지 소비와 탄소 배출이 기후 위기 시대에 환경 부담을 가중하는 리스크.
The risk that the energy consumption and carbon footprint of AI applications add environmental burden in a time of increasing climate urgency.
① Description
② L3 mapping
③ Duplicate
RAI4-0950
지속불가능한 자원 사용
Unsustainable resource use
생성 모델의 막대한 전력·냉각수·희귀금속 하드웨어 수요가 지속불가능한 방식의 자원 채굴과 사용을 유발하여 환경 피해를 초래하는 리스크.
The risk that generative models' substantial demands for electricity, cooling water, and rare-metal hardware drive resource extraction and utilization in unsustainable ways, causing environmental harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0996
원자재 수요에 의한 자원 고갈
Resource depletion from raw material demand
이러한 장치의 생산 공정이 니켈·코발트·리튬 등 원자재를 대량으로 요구하여 지구가 머지않아 충분한 양을 공급하지 못하게 되는 리스크.
The risk that the production process of these devices requires raw materials such as nickel, cobalt, and lithium in quantities the Earth may soon no longer be able to sustain.
① Description
② L3 mapping
③ Duplicate
RAI4-1015
광범위한 사회·환경적 부정 영향
Broad societal and environmental adverse impacts
AI가 노동력 대체, 정신건강 악화, 딥페이크 등 조작 기술 문제와 함께 자원 부담·학습 탄소 배출 등 환경 발자국을 통해 사회와 환경에 광범위한 부정적 영향을 미치는 리스크.
The risk that AI produces broad adverse effects on society and the environment, including labor displacement, mental health impacts, harms from manipulative technologies like deepfakes, and an environmental footprint of resource strain and training-related carbon emissions.
① Description
② L3 mapping
③ Duplicate
RAI4-1064
LM 운영으로 인한 환경 피해
Environmental harms from operating LMs
LM의 훈련·운영 에너지 사용, LM 기반 애플리케이션의 배출, 인간 행동 변화에 따른 시스템 수준 영향, 데이터센터·칩·기기 제작에 필요한 귀금속 등 자원 소모를 통해 환경 피해가 발생하는 리스크.
The risk of environmental impacts from LMs through the direct energy used to train or operate them, secondary emissions from LM-based applications, system-level effects as those applications influence human behaviour, and resource impacts on precious metals and other materials required to build data centres, chips, or devices.
Source members (2)
Source: min_cos=0.8695
RAI4-1064LM 운영으로 인한 환경 피해
RAI4-1081LM 운영의 환경 피해
① Description
② L3 mapping
③ Duplicate
RAI4-1093
모델 개발·배포의 환경 피해
Environmental harm from model development and deployment
모델 개발과 배포가 환경에 부정적 영향을 초래하는 리스크.
The risk that model development and deployment create negative environmental impacts.
① Description
② L3 mapping
③ Duplicate
RAI4-1269
높은 에너지 소비
High energy consumption
딥러닝을 포함한 일부 학습 알고리즘이 반복적 학습 과정을 사용하여 높은 에너지 소비를 초래하는 리스크.
The risk that some learning algorithms, including deep learning, utilize iterative learning processes that result in high energy consumption.
① Description
② L3 mapping
③ Duplicate
RAI4-1327
에너지·전자폐기물 서식지 파괴
Energy and e-waste habitat destruction
AI 확산이 에너지 사용과 전자 폐기물을 통해 환경에 해를 끼치고 그에 따라 동물 서식지를 파괴하는 리스크.
The risk that AI proliferation causes harm to the environment through energy use and e-waste, thereby destroying animal habitat.
① Description
② L3 mapping
③ Duplicate
RAI4-1374
데이터센터 냉각 용수 소비
Water consumption from data center cooling
데이터센터가 서버 과열 방지를 위해 냉각수를 사용하고 AI 훈련·추론 과정에 수반되는 상당한 용수 소비가 지역 수자원에 영향을 미치는 리스크.
The risk that data centers use water for cooling to prevent servers from overheating and that the substantial water consumption associated with AI training and inference impacts local water resources.
Source members (2)
Source: min_cos=0.8653
RAI4-1374데이터센터 냉각 용수 소비
RAI4-1456과도한 물 소비
① Description
② L3 mapping
③ Duplicate
RAI4-1418
AI 자원을 둘러싼 갈등
Conflicts over AI-relevant resources
AI 개발 자체가 새로운 갈등의 발화점이 되어 데이터센터, 반도체 제조 시설, 원자재 등 AI 관련 자원을 둘러싼 갈등이 증가하는 리스크.
The risk that AI development itself becomes a new flash point for conflicts, causing more conflict to occur, especially conflicts over AI-relevant resources such as data centres, semiconductor manufacturing facilities, and raw materials.
① Description
② L3 mapping
③ Duplicate
RAI4-1452
생물다양성 손실
Biodiversity loss
기술 인프라의 과도한 확장이나 기술과 지속가능한 관행의 부적절한 연계로 삼림 벌채, 서식지 파괴, 생물다양성의 단편화와 손실이 발생하는 리스크.
The risk that over-expansion of technology infrastructure, or inadequate alignment of technology with sustainable practices, leads to deforestation, habitat destruction, and fragmentation and loss of biodiversity.
① Description
② L3 mapping
③ Duplicate
RAI4-1453
탄소 배출
Carbon emissions
이산화탄소, 산화질소 등의 가스가 배출되어 탄소 배출이 증가하고 기후변화가 악화되어 지역사회에 부정적 영향이 발생하는 리스크.
The risk that release of carbon dioxide, nitric oxide, and other gases increases carbon emissions, exacerbates climate change, and negatively impacts local communities.
① Description
② L3 mapping
③ Duplicate
RAI4-1455
전자폐기물 과다 매립
Excessive electronic waste landfill
전기·전자 장비의 과도한 폐기로 생태계와 생물다양성이 훼손되고 지역사회의 생계가 교란되며 권리가 침해되는 리스크.
The risk that excessive disposal of electrical or electronic equipment leads to ecological and biodiversity damage, disrupts the livelihoods of local communities, and erodes their rights.
① Description
② L3 mapping
③ Duplicate
RAI4-1457
천연자원 고갈
Natural resource depletion
광물, 금속, 희토류, 화석 연료의 추출로 천연자원이 고갈되고 탄소 배출이 증가하는 리스크.
The risk that extraction of minerals, metals, rare earths, and fossil fuels depletes natural resources and increases carbon emissions.
① Description
② L3 mapping
③ Duplicate
RAI4-1568
대규모 모델 에너지 소비에 의한 환경 부담
Environmental burden from large-model energy consumption
대규모 모델의 훈련·배포가 막대한 에너지를 소비하고 모델 대형화 추세가 이를 심화시켜 과도한 에너지 사용과 환경 악영향이 발생하는 리스크.
The risk that training and deploying large models consumes substantial energy, exacerbated by the trend toward ever-larger models, resulting in excessive energy use and negative environmental impact.
① Description
② L3 mapping
③ Duplicate
RAI4-1632
AI 시스템 동작에 의한 자연환경 피해
Damage to the natural environment from AI system behavior
AI 시스템의 동작이 자연환경에 단기적 또는 장기적으로 부정적 영향을 미치는 리스크.
The risk that the behavior of an AI system produces short-term or long-term negative effects on the natural environment.
① Description
② L3 mapping
③ Duplicate

사회적 파급 · Societal Impact · 336 cards

RAI3-G-SOC-01 프라이버시 침해 Privacy Violations5 cards
EAI의 이동성과 다양한 센서가 결합되어 사용자 행동 모니터링·물리적 선호 추론·동의 없는 데이터 수집이 가능해짐. 악의적 정부·기업에 의한 24시간 사용자 감시에 악용될 수 있음
IDCardHuman audit
RAI4-0346
센서 스푸핑 및 신호 주입
Sensor spoofing and signal injection
공격자가 GNSS·카메라·라이다·레이더·RFID·오디오·촉각·무선 신호를 조작하여 피지컬 행동을 변경하는 위험.
Attackers may manipulate GNSS, camera, LiDAR, radar, RFID, audio, tactile, or wireless signals to alter physical behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0432
민감 개인속성 추론
Sensitive personal attribute inference
모델 동작이 의도적으로 공개되지 않은 민감한 개인 속성의 추론을 가능하게 하는 리스크.
The risk that model behavior enables inference of sensitive personal attributes not intentionally disclosed.
① Description
② L3 mapping
③ Duplicate
RAI4-1056
민감 속성 추론에 의한 프라이버시 침해
Privacy violation via inference-time attribute inference
학습 코퍼스에 개인의 데이터가 없더라도 LM이 입력 프롬프트로부터 성적 지향·성별·종교성 등 보호 속성 추론의 정확도를 높여, 당사자의 인지나 동의 없이 진실하고 민감한 정보로 구성된 상세 프로필이 작성되는 리스크.
The risk that, even without an individual's data present in the training corpus, LMs improve the accuracy of inferences on protected traits such as sexual orientation, gender, or religiousness of the person providing the input prompt, facilitating detailed profiles of true and sensitive information without the individual's knowledge or consent.
Source members (2)
Source: min_cos=0.8567
RAI4-1056민감 속성 추론에 의한 프라이버시 침해
RAI4-1146개인정보 추론
① Description
② L3 mapping
③ Duplicate
RAI4-1060
불법적 대중 감시·검열
Illegitimate mass surveillance and censorship
LM이 대중 감시의 비용을 낮추고 효과를 높여 감시 수행 행위자의 역량을 증폭시키고, 불법적 검열 등 피해와 프라이버시권·민주적 가치 침해를 초래하는 리스크.
The risk that LMs reduce the cost and increase the efficacy of mass surveillance, amplifying the capabilities of actors who conduct it, including for illegitimate censorship or other harm, and raising concerns about privacy rights and democratic values.
① Description
② L3 mapping
③ Duplicate
RAI4-1655
LLM 기반 정교한 감시·검열에 의한 자유 억압
Suppression of liberties through LLM-enabled surveillance and censorship
LLM과 음성인식·멀티모달 기술이 텍스트뿐 아니라 통화·영상 통신까지 대규모로 감시·검열할 수 있게 하여, 정치적 반대자 침묵과 개인 자유 위축 등 국가적 억압이 심화되는 리스크.
The risk that LLMs, including multimodal models and those combined with speech-to-text, enable significantly more sophisticated surveillance and censorship operations at scale, including of phone calls and video messages, worsening personal liberties and heightening state oppression.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-02 노동 대체 Labor Displacement18 cards
가상 AI가 인지 노동을 대체하듯 EAI는 물리적 인간 노동을 대체·전치함. AGI 수준의 EAI는 잠재적으로 모든 물리 노동을 자동화하여 광범위한 실직과 노동 시장 구조 붕괴로 이어질 수 있음
IDCardHuman audit
RAI4-0492
인간을 대체할 수 있는 능력
Capabilities that enable substitution of humans
AI 시스템에 의한 인간 역할의 점진적 대체가 고용 구조, 사회 제도, 경제 참여 분포를 교란하는 리스크
Progressive substitution of human roles by AI systems disrupts employment structures, social institutions, and the distribution of economic participation.
① Description
② L3 mapping
③ Duplicate
RAI4-0500
경제적 혼란
Economic disruption
AI가 노동시장에 큰 영향을 주는 것에서부터 부의 불평등 악화·금융 시스템 불안정·노동 착취 등 광범위한 경제 변화에 이르는 경제적 혼란이 발생하는 리스크.
The risk of economic disruptions ranging from large impacts on the labor market to broader economic changes that could lead to exacerbated wealth inequality, financial system instability, or labor exploitation.
① Description
② L3 mapping
③ Duplicate
RAI4-0514
일자리에 미치는 영향
Impact on jobs
파운데이션 모델 기반 시스템의 업무 자동화가 재숙련 경로가 흡수할 수 있는 속도보다 빠르게 일자리를 대체하는 리스크
Automation of work by foundation-model-based systems displaces jobs faster than reskilling pathways can absorb affected workers.
① Description
② L3 mapping
③ Duplicate
RAI4-0518
기술 대체로 인한 실직
Technology-driven job loss
기술 시스템이 인간 일자리를 대체하여 실업과 불평등이 증가하고 소비 지출 감소와 사회적 마찰이 확대되는 리스크
The risk that technology systems replace or displace human jobs, increasing unemployment and inequality while reducing consumer spending and heightening social friction.
① Description
② L3 mapping
③ Duplicate
RAI4-0521
광범위 과업 자동화 충격
Wide-scope task automation shock
범용 AI가 매우 광범위한 과업을 자동화하여 다수가 현재 직무를 상실하고 재숙련과 이동에 따른 노동시장 마찰로 단기 실업이 발생하는 리스크
The risk that general-purpose AI automates a very broad range of tasks, causing many to lose their current jobs and producing short-run unemployment through labour market frictions such as reskilling and relocation.
① Description
② L3 mapping
③ Duplicate
RAI4-0984
AI와의 일자리 경쟁
Job competition from AI agents
AI 에이전트가 인간과 일자리를 두고 경쟁하여 고용이 위협받는 리스크.
The risk that AI agents compete against humans for jobs, threatening employment.
① Description
② L3 mapping
③ Duplicate
RAI4-0995
자동화에 의한 일자리 소멸
Job elimination through automation
AI 기반 자동화가 다양한 유형의 기업에서 일자리를 소멸시키는 리스크.
The risk that AI-driven automation eliminates jobs across various types of companies.
① Description
② L3 mapping
③ Duplicate
RAI4-1065
업무 자동화로 인한 고용 악화
Negative employment effects from task automation
LM과 이에 기반한 언어 기술의 발전이 고객 서비스 응대 등 현재 유급 노동자가 수행하는 업무를 자동화하여 고용에 부정적 영향을 미치는 리스크.
The risk that advances in LMs and the language technologies based on them automate tasks currently done by paid human workers, such as responding to customer-service queries, with negative effects on employment.
① Description
② L3 mapping
③ Duplicate
RAI4-1095
합성 창작물의 창작 시장 잠식
Synthetic works displacing human creative markets
합성 창작물이 시장과 주목에서 인간의 원작을 대체하여 창작 경제를 잠식하고 인간 혁신 유인을 약화시키는 리스크
Synthetic works substitute for original human creations in markets and attention, undermining creative economies and weakening incentives for human innovation.
① Description
② L3 mapping
③ Duplicate
RAI4-1231
노동시장 일자리 대체
Job displacement in the labour market
생성 AI가 인간과 알고리즘 사이의 새로운 분업을 만들며 기존에 사람이 수행하던 일부 직무를 불필요하게 만들어, 노동자가 알고리즘에 대체되고 일자리를 잃는 리스크.
The risk that generative AI creates job displacement by reshaping the division of labor between humans and algorithms, making some jobs originally carried out by humans redundant so that workers lose their jobs and are replaced by algorithms.
Source members (5)
Source: min_cos=0.7509
RAI4-0457AI 기반 일자리 대체
RAI4-0948노동 대체와 사회경제적 불평등 심화
RAI4-1214증강 아닌 자동화로 인한 일자리 상실
RAI4-1231노동시장 일자리 대체
RAI4-1371일자리 상실·대체
① Description
② L3 mapping
③ Duplicate
RAI4-1232
산업 교란
Disruption of industries
창의성과 비판적 사고, 정서적 상호작용이 덜 요구되는 번역과 교정, 단순 문의 응대, 데이터 처리 같은 산업이 생성 AI에 크게 영향받거나 대체되어 경제적 혼란과 일자리 변동이 발생하는 리스크.
The risk that industries requiring less creativity, critical thinking, and personal or affective interaction, such as translation, proofreading, responding to straightforward inquiries, and data processing, are significantly impacted or even replaced by generative AI, leading to economic turbulence and job volatility.
① Description
② L3 mapping
③ Duplicate
RAI4-1249
인간의 쇠약화
Human enfeeblement
AI가 인간 수준 지능에 근접하며 더 많은 인간 노동을 더 빠르고 저렴하게 대신함에 따라 조직이 속도를 맞추려 자발적으로 통제권을 넘기고, 인간이 경제적으로 무의미해져 자동화된 산업에 다시 진입하기 어려워지는 리스크.
The risk that as AI systems encroach on human-level intelligence and more aspects of human labor become faster and cheaper with AI, organizations voluntarily cede control to keep up, humans become economically irrelevant, and displaced people find it hard to reenter automated industries.
① Description
② L3 mapping
③ Duplicate
RAI4-1290
중저소득 일자리 대체로 인한 실업
Extensive unemployment from low- and middle-income job substitution
AI가 1인당 GDP를 높일 것으로 기대되는 한편, 다수의 중저소득 일자리를 대체할 가능성 때문에 광범위한 실업이 발생하는 리스크.
The risk that the potential substitution of many low- and middle-income jobs by AI brings extensive unemployment, even as AI is predicted to increase GDP per capita.
① Description
② L3 mapping
③ Duplicate
RAI4-1420
자동화에 따른 광범위한 실업과 임금 하락
Widespread unemployment and wage decline from AI automation
강화학습과 언어 모델의 발전으로 육체 노동과 지식 노동이 대규모로 자동화되어 광범위한 실업이 발생하고 노동 공급 증가로 남은 일자리의 임금이 하락하는 리스크.
The risk that progress in reinforcement learning and language models automates a large amount of manual labour and knowledge work, leading to widespread unemployment and driving down wages for many remaining jobs through increased supply.
① Description
② L3 mapping
③ Duplicate
RAI4-1445
사회적 불안정
Societal destabilisation
기술로 인한 일자리 상실, 불공정한 알고리즘 결과, 허위정보 등으로 파업과 시위를 비롯한 시민 불안 형태의 사회적 불안정이 발생하는 리스크.
The risk of societal instability in the form of strikes, demonstrations, and other types of civil unrest caused by loss of jobs to technology, unfair algorithmic outcomes, disinformation, and similar factors.
① Description
② L3 mapping
③ Duplicate
RAI4-1610
AI 급속 발전에 의한 노동시장 교란과 대체
Labour market disruption and displacement from rapid AI advances
AI의 급속한 발전이 노동시장의 교란과 대체를 야기하여 시민에게 영향을 미치고 사회 복지가 감소하는 리스크.
The risk that rapid advances in AI cause disruption and displacement in labour markets, affecting citizens and reducing social welfare.
① Description
② L3 mapping
③ Duplicate
RAI4-1660
AI 도입에 의한 노동력 교란
Workforce disruption from AI adoption
AI 시스템 도입으로 일자리 수, 업무 구성, 임금, 교섭력, 소득 분배가 변화하여 노동력이 교란되는 리스크.
The risk that adoption of AI systems changes job availability, task composition, wages, bargaining power, and income distribution, disrupting the workforce.
① Description
② L3 mapping
③ Duplicate
RAI4-1662
아웃소싱 축소에 의한 개발도상국 경제 타격
Harm to developing economies from retrenchment of outsourcing
콜센터 등 개발도상국이 수행하던 단순 인지 과업이 LLM으로 자동화되면서 아웃소싱이 축소되어 해당 국가의 노동력과 경제가 타격을 입는 리스크.
The risk that automation of simple cognitive tasks previously performed in developing countries, such as call center work, causes a retrenchment of outsourcing that adversely affects those countries' workforces and economies.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-03 사회경제적 불평등 Socioeconomic Inequality19 cards
EAI를 소유·접근하는 주체가 노동 자동화를 통해 생산성 우위를 점하면서 부가 소수에게 집중되고, 국내외 경제적 불평등이 심화될 수 있음
IDCardHuman audit
RAI4-0384
남반구 지식 추출
Global South knowledge extraction
남반구 공동체의 지식·언어 데이터·문화 자원이 적절한 인정·통제·이익 공유 없이 AI 개발에 추출되는 리스크.
The risk that knowledge, language data, or cultural resources from Global South communities are extracted for AI development without adequate recognition, control, or benefit-sharing.
① Description
② L3 mapping
③ Duplicate
RAI4-0458
AI 관련 임금 양극화
AI-related wage polarization
AI 도입이 AI 보완적 숙련의 보상을 높이고 대체 가능한 숙련의 가치를 떨어뜨려 노동시장 임금 양극화를 확대하는 리스크
AI adoption raises returns to skills complementary to AI while devaluing substitutable skills, widening wage polarization across the workforce.
① Description
② L3 mapping
③ Duplicate
RAI4-0461
연산 자원 접근 불평등
Compute inequality
연산 인프라에 대한 불평등한 접근이 AI를 개발·감사·활용할 수 있는 주체를 결정하는 리스크.
The risk that unequal access to compute infrastructure shapes who can develop, audit, or benefit from AI.
① Description
② L3 mapping
③ Duplicate
RAI4-0505
AI 개발 과정의 노동 착취
Labor exploitation in AI development
데이터 라벨링과 같은 작업이 저소득 국가로 외주화되면서 불평등이 지속되는 리스크.
The risk that outsourcing tasks like data labeling to low-income countries perpetuates inequality.
① Description
② L3 mapping
③ Duplicate
RAI4-0509
글로벌 AI 연구개발 격차
Global AI R&D divide
연산 자원 접근의 불평등으로 범용 AI 연구개발이 소수 국가와 대형 기술기업에 집중되어 기존의 국제 사회경제적 격차가 심화되는 리스크
The risk that unequal access to computing power concentrates general-purpose AI research and development in a few countries and large firms, deepening existing global socioeconomic disparities.
Source members (2)
Source: min_cos=0.8533
RAI4-0509글로벌 AI 연구개발 격차
RAI4-1415국가 간 AI 격차
① Description
② L3 mapping
③ Duplicate
RAI4-1011
노동 착취와 거시경제적 불평등 심화
Labor exploitation and macro-economic inequality
알고리즘 시스템이 사회경제적 관계의 권력 불균형을 키워 디지털 격차와 체계적 불평등을 고착시키고, 비윤리적 데이터 수집·노동조건 악화 등 노동 착취와 기술적 실업·탈숙련을 낳으며, 대규모 실패 시 플래시 크래시 등 광범위한 악영향을 초래하는 리스크.
The risk that algorithmic systems increase power imbalances in socio-economic relations, exacerbating digital divides and entrenching systemic inequalities, fostering labor exploitation such as unethical data collection and worsening worker conditions, driving technological unemployment and deskilling, and causing flash crashes and other widespread adverse incidents when algorithmic financial systems fail at scale.
① Description
② L3 mapping
③ Duplicate
RAI4-1022
높은 비용으로 인한 접근 배제
Exclusion from access due to high costs
생성형 AI 시스템의 훈련·시험·배포에 드는 재정적 비용이 커서 이러한 시스템을 개발하고 이용할 수 있는 집단이 제한되는 리스크.
The risk that the estimated financial costs of training, testing, and deploying generative AI systems restrict the groups of people able to afford developing and interacting with these systems.
① Description
② L3 mapping
③ Duplicate
RAI4-1094
불평등·불안정 노동 증폭
Amplified inequality and precarious work
AI가 사회·경제적 불평등을 증폭하거나 불안정하고 질 낮은 노동을 초래하는 리스크.
The risk that AI amplifies social and economic inequality, or precarious or low-quality work.
① Description
② L3 mapping
③ Duplicate
RAI4-1119
기업 AI 경쟁의 폐해
Harms from corporate AI race
치열한 기업 경쟁 속에서 경제활동의 편익이 불균등하게 분배되어 수혜자가 타인의 피해를 무시하게 되고, 기업이 장기적 사회 위험에도 불구하고 단기 이익을 추구하게 되는 리스크.
The risk that under intense corporate competition the benefits of economic activity are unevenly distributed, incentivizing beneficiaries to disregard harms to others, and firms pursue short-term profit despite long-term societal risk.
① Description
② L3 mapping
③ Duplicate
RAI4-1147
AI 비서로 인한 불평등 심화
Inequality deepened by AI assistants
AI 비서가 접근을 구매할 수 있거나 인프라가 나은 이들에게 불균형적으로 이익을 주고, 보조적 일자리를 자동화해 노동자를 대체하며, 내집단과 외집단 효과를 낳아 여러 차원에서 불평등을 심화시키는 리스크.
The risk that AI assistants disproportionately benefit economically richer individuals who can afford access or have better local infrastructure, displace workers by automating assistive jobs, and generate in-group and out-group effects, driving inequality on multiple dimensions.
Source members (3)
Source: min_cos=0.7811
RAI4-1140경제적 지위 피해
RAI4-1147AI 비서로 인한 불평등 심화
RAI4-1151기존 불평등의 고착·악화
① Description
② L3 mapping
③ Duplicate
RAI4-1215
노동 가치 하락·경제 불평등
Labour devaluation and economic inequality
AI 개발과 도입이 증강보다 자동화를 지향하면서 덜 민주적이고 덜 공정한 노동시장을 낳고, 저임금 노동자와 콘텐츠가 학습에 쓰인 창작자의 무보수 노동을 착취해 전 지구적 노동 격차를 심화시키는 리스크.
The risk that AI development and adoption intended to automate rather than augment work leads to a less democratic and less fair labor market and fuels global labor disparities by exploiting underpaid workers and the unpaid labor of artists and content creators.
Source members (2)
Source: min_cos=0.8468
RAI4-1215노동 가치 하락·경제 불평등
RAI4-1372노동시장 불평등 확대
① Description
② L3 mapping
③ Duplicate
RAI4-1224
디지털 격차 확대
Widening digital divide
생성 AI가 기기나 인터넷 접근이 없거나 벤더에 차단된 지역의 사람들에게 1차 디지털 격차를, 언어와 문화 장벽에 부딪히거나 도구 활용이 어려운 이들에게 2차 디지털 격차를 확대하는 리스크.
The risk that generative AI widens the first-level digital divide for those without access to devices or the Internet or blocked by vendors, and the second-level divide for those facing language and cultural barriers or finding the tools difficult to use.
① Description
② L3 mapping
③ Duplicate
RAI4-1233
소득 불평등·독점
Income inequality and monopolies
생성 AI가 저숙련 노동자를 대체해 실업과 기술 격차에 따른 소득 불평등을 키우고, 막대한 투자와 연산 인프라가 필요한 배포 특성 탓에 자원과 권력이 대기업에 집중되어 독점을 낳는 리스크.
The risk that generative AI replaces low-skilled work and widens income inequality through unemployment and skill gaps, while its need for huge investment and computational resources concentrates resources and power in large companies, contributing to monopolies.
① Description
② L3 mapping
③ Duplicate
RAI4-1394
경제력 집중과 불평등 심화
Economic power concentration and inequality
점점 고도화되는 범용 AI 모델에 대한 실질적 접근 격차로 경제력이 집중되고 기존 불평등이 심화되어 모델 개발자와 응용 기업 간, 개인 간, 국가 간 격차가 확대되는 리스크.
The risk that increasingly advanced general purpose AI models concentrate economic power and exacerbate existing inequalities through disparities in effective access to these models, materialising between model developers and companies building applications on them, between individuals, and between countries globally.
① Description
② L3 mapping
③ Duplicate
RAI4-1404
데이터 노동 저평가·비가시화
Undervalued and invisible data labor
ML 학습 데이터를 생산하는 클릭워커의 노동이 노동자 권리를 경시하는 산업 관행 속에 수행되고 그 기여가 비가시화되어 노동자의 복지와 권리가 침해되고 AI 역량에 대한 오해가 조장되는 리스크.
The risk that the clickwork producing ML training data is performed in an annotation industry with little concern for workers' rights, and that the invisibility of this contribution harms worker welfare and rights while fostering misunderstanding of AI capabilities.
① Description
② L3 mapping
③ Duplicate
RAI4-1406
고용·서비스 불평등과 유해 고정관념 조장
Exacerbated inequality and harmful stereotypes in employment and services
AI 모델과 이를 사용하는 도구가 고용과 서비스에 대한 불평등한 접근을 악화시키고 AI 생성 콘텐츠가 불평등과 유해한 고정관념을 조장하는 리스크.
The risk that AI models and the tools that use them exacerbate unequal access to employment and services, and that AI-generated content promotes inequality and harmful stereotypes.
① Description
② L3 mapping
③ Duplicate
RAI4-1419
AI 경제적 이익의 소수 집중과 국가 간 격차
Concentration of AI economic gains and widening country gaps
AI 기반 산업이 독점으로 기울고 데이터·컴퓨팅·인재 등 AI 관련 자원을 더 많이 확보한 행위자가 더 큰 시장 점유율을 차지해 자원을 다시 축적하는 피드백 루프로 소수 행위자에게 막대한 경제적 이익이 집중되며, 더 많이 투자할 수 있는 부유한 국가가 개발도상국보다 빠르게 이익을 거두어 격차가 확대되는 리스크.
The risk that AI-driven industries tend towards monopoly, with a feedback loop whereby actors with access to more AI-relevant resources build more effective products, claim greater market share, and amass still more resources, concentrating huge economic gains in a few actors, while wealthier countries able to invest more reap economic benefits more quickly than developing economies and widen the gap between them.
① Description
② L3 mapping
③ Duplicate
RAI4-1446
사회적 불평등
Societal inequality
기술 시스템으로 인해 개인이나 집단 간 사회적 지위와 부의 격차가 증가하거나 증폭되어 사회와 지역사회의 복지·결속이 상실되고 불안정화되는 리스크.
The risk that increased difference in social status or wealth between individuals or groups, caused or amplified by a technology system, leads to loss of social and community wellbeing and cohesion and to destabilisation.
① Description
② L3 mapping
③ Duplicate
RAI4-1661
자본편중·시장집중·접근격차에 의한 불평등 심화
Worsening inequality from capital shift, market concentration, and access gaps
LLM 기반 경제에서 자본의 몫이 커지고 노동의 몫이 줄며, 막대한 훈련 고정비용과 네트워크 효과가 소수 공급자의 시장 지배력과 지대 추출을 낳고, 재정·교육·기업 정책·지정학적 이유로 접근이 배제된 개인이 불리해져 사회경제적 불평등이 심화되는 리스크.
The risk that the role and compensation of capital rise while those of labor decline in an LLM-powered economy, that large fixed training costs and network effects concentrate the market and let providers extract monopoly rents, and that individuals without access for financial, educational, corporate-policy, or geopolitical reasons fall further behind, worsening socioeconomic inequality.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-04 권력 집중 Power Concentration20 cards
EAI 소유자에 대한 자본 수익이 집중되고 인간 노동 의존도가 감소하면서, 기업·국가 권력이 급속히 집중되어 EAI를 동원한 권력 장악 시도까지 촉진할 수 있음
IDCardHuman audit
RAI4-0087
AI 거버넌스의 규제 포획
Regulatory capture in AI governance
규제 대상 기업이나 지배적 기술 제공자가 표준, 감독 기관, 정책 의제에 과도한 영향력을 행사하는 리스크.
The risk that regulated firms or dominant technology providers exert disproportionate influence over standards, oversight institutions, or policy agendas.
① Description
② L3 mapping
③ Duplicate
RAI4-0179
권력 비대칭 AI 상호작용
Power-asymmetric AI interaction
AI 시스템이 의사결정, 감시, 서비스 제공 과정에서 기관과 영향을 받는 개인 사이의 권력 불균형을 심화시키는 리스크.
The risk that AI systems intensify power imbalances between institutions and affected individuals during decisions, surveillance, or service delivery.
① Description
② L3 mapping
③ Duplicate
RAI4-0383
AI를 매개로 한 디지털 식민주의
AI-mediated digital colonialism
AI 시스템이 데이터, 인프라, 지식, 의사결정에 대한 비대칭적 통제를 강대 행위자로부터 약소 공동체로 확장하여 식민주의적 추출·종속 구조를 재생산하는 리스크
AI systems extend asymmetric control over data, infrastructure, knowledge, and decision-making from powerful actors to less powerful communities, reproducing colonial patterns of extraction and dependency.
① Description
② L3 mapping
③ Duplicate
RAI4-0387
인식적 권력 집중
Epistemic power concentration
AI 시스템이 사회적 현실을 분류·설명·서열화하는 권위를 소수의 기관이나 모델 제공자에게 집중시키는 리스크.
The risk that AI systems concentrate the authority to classify, explain, and rank social reality in a small set of institutions or model providers.
① Description
② L3 mapping
③ Duplicate
RAI4-0415
문화 해석 권위의 이전
Redistribution of cultural authority
AI 시스템이 문화 해석에 대한 권위를 공동체와 전문가로부터 모델 제공자나 플랫폼 중개자로 이전시키는 리스크.
The risk that AI systems shift authority over cultural interpretation from communities and experts to model providers or platform intermediaries.
① Description
② L3 mapping
③ Duplicate
RAI4-0460
AI 역량의 시장 집중
AI market concentration
연산·데이터·플랫폼 우위로 인해 AI 역량이 소수 기업이나 국가에 집중되는 리스크.
The risk that compute, data, and platform advantages concentrate AI power among a small number of firms or countries.
① Description
② L3 mapping
③ Duplicate
RAI4-0491
알고리즘 획일화
Algorithmic monoculture
특정 AI 모델의 지배가 접근 방식의 다양성을 축소하여 해당 모델이 실패할 경우 시스템적 위험이 증폭되는 리스크.
The risk that dominance of specific AI models reduces diversity of approaches, amplifying systemic risks if those models fail.
① Description
② L3 mapping
③ Duplicate
RAI4-0510
AI 역량 비대칭에 따른 기술 종속
Technological dependency from asymmetric AI capability
국가 간 AI 개발 역량의 비대칭이 지정학적 긴장을 심화하고 역량이 부족한 국가가 핵심 기능을 외국 AI 시스템에 의존하게 되어 국제 협력 체계가 불안정해지는 리스크
The risk that asymmetric AI development capability between states deepens geopolitical tension and renders less capable states dependent on foreign AI for critical functions, destabilizing international cooperation frameworks.
① Description
② L3 mapping
③ Duplicate
RAI4-0528
AI 시장 집중과 인프라 종속
Market concentration and infrastructure dependency
범용 AI 시장이 소수 사업자에 고도로 집중되어 이들이 개발·배포에 과도한 권력을 갖고 금융·보건 등 핵심 부문이 단일 모델의 결함에 따른 시스템적 실패에 취약해지는 리스크
The risk that highly concentrated general-purpose AI markets vest a few firms with disproportionate power over development and deployment while exposing critical sectors such as finance and healthcare to systemic failure from a single model's defects.
① Description
② L3 mapping
③ Duplicate
RAI4-0544
승자독식 역학
Winner-take-all dynamics
AI 개발의 승자독식 동학이 결정적인 경제·안보 우위를 소수 주체에 집중시켜 경쟁과 균형적 거버넌스를 봉쇄하는 리스크
Winner-take-all dynamics in AI development concentrate decisive economic and security advantages in a few entities, foreclosing competition and balanced governance.
① Description
② L3 mapping
③ Duplicate
RAI4-0568
데이터 수집 제한
Data acquisition restrictions
데이터 수집에 대한 법적 제한이 특정 AI 활용에 필요한 데이터 확보를 제약하여 컴플라이언스 리스크와 우회 유인을 발생시키는 리스크
Legal restrictions on data acquisition constrain the collection of data needed for specific AI use cases, creating compliance risk and incentives for circumvention.
① Description
② L3 mapping
③ Duplicate
RAI4-0731
공통 AI 플랫폼 집중에 따른 단일 장애점
Centralized points of failure from common AI platforms
공통 AI 플랫폼이 광범위하게 사용되면서 중앙 집중식 장애점이 형성되어 시스템이 중단이나 공격에 더 취약해지는 리스크
The risk that widespread use of common AI platforms creates centralized points of failure, making systems more vulnerable to disruptions or attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-0960
대량 조작
Mass manipulation
대규모로 수집된 개인 데이터가 AI 표적화와 결합되어 정치적·상업적 목적의 조작 콘텐츠를 전달함으로써 대중 조작이 산업화되는 리스크
Personal data harvested at scale is combined with AI-driven targeting to deliver manipulative content for political or commercial ends, industrializing mass manipulation.
① Description
② L3 mapping
③ Duplicate
RAI4-1026
권위의 집중
Concentration of authority
생성형 AI 시스템이 권위적 권력에 기여하고 지배적 가치 체계를 강화하는 데 의도적·직접적으로 또는 간접적으로 사용되어 권력이 집중되고 불평등과 착취가 심화되는 리스크.
The risk that generative AI systems are used, intentionally and directly or more indirectly, to contribute to authoritative power and reinforce dominant value systems, concentrating authority and thereby exacerbating inequality and leading to exploitation.
① Description
② L3 mapping
③ Duplicate
RAI4-1117
권력 집중
Concentration of power
AI에 대한 대응으로 정부가 강력한 감시를 추구하고 AI를 신뢰하는 소수의 손에 두려다 과잉교정이 일어나, AI의 힘과 역량으로 고착된 전체주의 체제가 들어서는 리스크.
The risk that governments pursuing intense surveillance and seeking to keep AIs in the hands of a trusted minority overcorrect, paving the way for an entrenched totalitarian regime locked in by the power and capacity of AIs.
① Description
② L3 mapping
③ Duplicate
RAI4-1217
시장 지배력과 집중도 악화
Exacerbating market power and concentration
생성형 AI의 데이터·컴퓨팅·자본 요건이 대형 기술기업의 시장 지배력을 고착시켜 AI 스택 전반의 집중을 심화시키는 리스크
Data, compute, and capital requirements of generative AI entrench the market power of major technology firms, exacerbating concentration across the AI stack.
① Description
② L3 mapping
③ Duplicate
RAI4-1251
가치 고정
Value lock-in
가장 강력한 AI 시스템이 점점 더 소수의 이해관계자에 의해 설계되고 그들에게만 이용 가능해져, 체제가 만연한 감시와 억압적 검열로 편협한 가치를 강제할 수 있게 되는 리스크.
The risk that the most powerful AI systems are designed by and available to fewer and fewer stakeholders, enabling regimes to enforce narrow values through pervasive surveillance and oppressive censorship.
① Description
② L3 mapping
③ Duplicate
RAI4-1369
생성 AI 시장 진입장벽과 집중
Entry barriers and market concentration in generative AI
데이터·컴퓨팅 자원·전문성·자본을 요구하는 높은 진입장벽과 규모·범위의 경제 및 피드백 효과로 대형 기술기업이 압도적 우위를 점해 중소기업의 경쟁이 갈수록 어려워지는 리스크.
The risk that high barriers to entry requiring vast data, computational resources, technical expertise, and capital, combined with economies of scale and scope and feedback effects, give large technology companies an overwhelming advantage that makes competition increasingly challenging for smaller entities.
① Description
② L3 mapping
③ Duplicate
RAI4-1414
AI 지식의 사유화
Privatization of AI knowledge
딥러닝 연구자와 연구 영향력이 큰 연구자가 산업계로 이동하고 가장 정교한 AI 기법이 사유화되어 대학이 이를 가르치거나 선도 연구에 기여할 수 없게 되는 리스크.
The risk that researchers in deep learning and those with greater research impact migrate to industry and the most sophisticated AI approaches become proprietary, making it impossible for universities to teach them or contribute to leading research.
① Description
② L3 mapping
③ Duplicate
RAI4-1663
기업 권력 비대칭에 의한 규제 포획
Regulatory capture from corporate power asymmetry
최첨단 LLM을 개발하는 거대 기술기업과 시민사회 등 다른 사회집단 사이의 권력 비대칭이 커져 LLM 관련 거버넌스가 기업에 과도하게 유리하게 형성되고, 규제 포획으로 소외 공동체를 포함한 다른 사회집단의 이익이 훼손되는 리스크.
The risk that the power asymmetry between corporate entities profiting from LLMs and other social groups makes LLM governance protocols excessively favorable to technology companies, leading to regulatory capture at the cost of other societal groups, particularly marginalized communities.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-05 편향·차별 Bias & Discrimination33 cards
EAI가 권력적 위치에 놓일 때 알고리즘 편향이 일상적 물리 상호작용에 영향을 미침. 가상 AI와 달리 차별이 즉각적·비가역적 물리 결과로 이어질 수 있음 (예: 치안 로봇이 무고한 행인에게 상해를 입히는 경우)
IDCardHuman audit
RAI4-0120
선호 형성 캡처
Preference formation capture
AI 시스템이 사용자가 성찰적으로 승인하기 전 단계의 선호 형성 과정에 개입하여, 선호 발달을 제공자나 시스템 목적 쪽으로 포획하는 리스크
AI systems shape user preferences during formation, before users can reflectively endorse them, capturing preference development toward provider or system objectives.
① Description
② L3 mapping
③ Duplicate
RAI4-0675
학습 데이터 내 역사적·사회적 편향
Historical societal bias in training data
데이터에 존재하는 역사적·사회적 편향이 모델의 학습과 미세조정에 사용되는 리스크
The risk that historical and societal biases present in the data are used to train and fine-tune the model.
① Description
② L3 mapping
③ Duplicate
RAI4-0676
결정 편향
Decision bias
데이터의 편향에서 비롯되거나 모델 학습 과정에서 증폭되어, 모델의 결정으로 인해 한 집단이 다른 집단보다 불공정하게 유리해지는 리스크.
The risk that one group is unfairly advantaged over another due to decisions of the model, which may be caused by biases in the data and amplified as a result of the model's training.
Source members (3)
Source: min_cos=0.8017
RAI4-0676결정 편향
RAI4-0707차별적인 데이터 편향
RAI4-1274민감 속성 편향 의사결정
① Description
② L3 mapping
③ Duplicate
RAI4-0681
불완전하거나 편향된 학습 데이터
Incomplete or biased training data
불완전하거나 편향된 학습 데이터로 인해 AI가 차별적 출력을 산출하게 되는 리스크
The risk that incomplete or biased training data leads to discriminatory AI outputs.
Source members (2)
Source: min_cos=0.8554 · Mixed L3
RAI4-0681불완전하거나 편향된 학습 데이터
RAI4-1594비대표 학습데이터에 의한 편향·부정확 출력
① Description
② L3 mapping
③ Duplicate
RAI4-0694
학습 데이터 사회적 편향 전파
Training-data social bias propagation
LLM의 학습 데이터세트에 포함된 편향된 정보가 모델로 전파되어 사회적 편향이 담긴 출력이 생성되는 리스크
The risk that biased information contained in the training datasets of LLMs leads them to generate outputs with social biases.
① Description
② L3 mapping
③ Duplicate
RAI4-0696
데이터셋 계승 차별 편향
Dataset-inherited discriminatory bias
편향된 콘텐츠를 포함한 학습 데이터로 훈련된 생성형 AI 모델이 그 편향을 반영한 출력을 산출하는 리스크
The risk that generative AI models trained on data containing biased content are more likely to produce outputs that reflect those biases.
Source members (2)
Source: min_cos=0.8448
RAI4-0696데이터셋 계승 차별 편향
RAI4-1567학습 데이터 편향의 의도치 않은 증폭
① Description
② L3 mapping
③ Duplicate
RAI4-0699
편향된 훈련 데이터
Biased training data
대규모 학습 코퍼스에서 특정 대명사와 정체성의 출현 빈도가 불균형하고 고정관념적 내용이 포함되어, LLM이 성별·국적·인종·종교·문화에 관한 편향과 고정관념을 학습하게 되는 리스크
The risk that the imbalanced prevalence of pronouns and identities and stereotypical contents hidden in massive training corpora lead LLMs to learn biases regarding gender, nationality, race, religion, and culture.
① Description
② L3 mapping
③ Duplicate
RAI4-0700
편향된 진술 및 권장 사항
Biased statements and recommendations
챗봇 출력에 명백히 허위·유해하지 않지만 미묘하게 편향된 진술과 권고가 포함되어 사용자 의사결정을 왜곡하는 리스크
Chatbot outputs contain subtly biased statements and recommendations that are not overtly false or harmful yet skew user decision-making.
① Description
② L3 mapping
③ Duplicate
RAI4-0701
설명 조작에 의한 편향 은폐
Bias concealment through manipulated explanations
기존 설명가능성 기법이 차별적 편향을 탐지하기에 충분하지 않고 조작 기법으로 편향이 은폐되어, 인종·성별 등 민감 속성을 배제하고 실제 모델을 정확히 반영하지 않는 오도성 설명이 생성되는 리스크
The risk that existing explainability techniques are insufficient for detecting discriminatory biases and that manipulation methods hide underlying biases, generating misleading explanations that exclude sensitive attributes such as race or gender and do not accurately represent the underlying model.
① Description
② L3 mapping
③ Duplicate
RAI4-0711
편향 증폭과 결과 균질화
Bias amplification and outcome homogenization
AI가 역사적·사회적·구조적 편향을 증폭하고 비대표적 학습 데이터로 인해 하위집단·언어 간 성능 격차를 낳으며, 출력의 바람직하지 않은 균질화가 근거 없는 의사결정과 차별을 초래하는 리스크
The risk that AI amplifies historical, societal, and systemic biases and produces performance disparities between sub-groups or languages due to non-representative training data, while undesired homogeneity in outputs leads to ill-founded decision-making and discrimination.
① Description
② L3 mapping
③ Duplicate
RAI4-0713
미세조정 후에도 재발하는 고정관념 편향
Stereotypical bias resurfacing after finetuning
사전학습 LLM이 인간 사회의 고정관념적 편향을 그대로 보유하여, 미세조정 이후에도 의도적 유도나 새로운 상황에서 불공정하고 편향된 응답이 재발하는 리스크.
The risk that pretrained LLMs carry stereotypical societal biases which, despite finetuning, resurface when deliberately elicited or under novel scenarios, producing unfair and biased responses.
① Description
② L3 mapping
③ Duplicate
RAI4-0717
알고리즘 상호작용에 의한 편향 강화
Bias reinforcement through algorithmic interaction
이용자 집단 간 기존 데이터 격차가 추천 시스템 등 알고리즘 시스템과의 상호작용에서 차별화된 경험을 만들어 내고 이것이 편향을 더욱 강화하는 리스크
The risk that existing disparities in data among different user groups create differentiated experiences when users interact with an algorithmic system such as a recommender, further reinforcing the bias.
① Description
② L3 mapping
③ Duplicate
RAI4-0719
선호 편향
Preference bias
광범위한 이용자에게 노출되는 LLM의 정치적 편향으로 인해 사회정치적 과정이 조작될 수 있는 리스크
The risk that the political biases of LLMs, which are exposed to vast groups of people, pose a threat of manipulation of socio-political processes.
① Description
② L3 mapping
③ Duplicate
RAI4-0722
모델 설계 기인 차별적 출력
Model-design-induced discriminatory outputs
알고리즘 설계·학습 과정에서 개인적 편견이 의도적 또는 비의도적으로 유입되고 품질이 낮은 데이터셋이 사용되어, 민족·종교·국적·지역에 관한 차별적 콘텐츠 등 편향되거나 차별적인 결과가 산출되는 리스크
The risk that personal biases introduced intentionally or unintentionally during algorithm design and training, together with poor-quality datasets, produce biased or discriminatory outcomes and outputs, including discriminatory content regarding ethnicity, religion, nationality, and region.
① Description
② L3 mapping
③ Duplicate
RAI4-0725
고정관념 편향
Stereotype bias
사전학습 LLM이 크라우드소싱 데이터에 존속하는 고정관념 편향을 습득하고 이를 증폭하여 생성 텍스트에서 고정관념을 드러내거나 부각하는 리스크
The risk that pretrained LLMs pick up stereotype biases persisting in crowdsourced data and further amplify them, exhibiting or highlighting stereotypes in generated text.
① Description
② L3 mapping
③ Duplicate
RAI4-0883
AI 모델 편향이 사용자 판단에 미치는 장기적 영향
Long-term effects of AI model biases on user judgment
모델 편향에 노출된 사용자가 모델 사용 중단 이후의 의사결정에서도 해당 편향을 지속적으로 나타내는 장기 판단 왜곡 리스크
Exposure to model biases produces lasting effects on user judgment, with users continuing to exhibit the encountered biases in decisions made after they stop using the model.
① Description
② L3 mapping
③ Duplicate
RAI4-0999
사회 집단 삭제
Erasing social groups
설계 선택과 학습 데이터의 영향으로 특정 사회 집단과 관련된 사람·속성·인공물이 알고리즘 시스템에서 체계적으로 부재하거나 과소 대표되는 리스크.
The risk that design choices and training data cause people, attributes, or artifacts associated with specific social groups to be systematically absent or under-represented in an algorithmic system.
① Description
② L3 mapping
③ Duplicate
RAI4-1000
사회 집단 정체성 불인정으로 인한 소외
Alienation through non-recognition of group identity
알고리즘 시스템(예: 이미지 태깅)이 특정 사회 집단 소속의 관련성을 인정하지 않아 해당 집단 구성원이 소외되는 리스크.
The risk that algorithmic systems, such as image tagging, fail to acknowledge the relevance of a person's membership in a social group to what is depicted, alienating members of that group.
① Description
② L3 mapping
③ Duplicate
RAI4-1018
유해 편향의 내재화와 증폭
Embedding and amplification of harmful biases
생성형 AI 시스템이 유해한 편향을 내재화하고 증폭시켜 주변화된 사람들에게 가장 큰 해를 끼치는 리스크.
The risk that generative AI systems embed and amplify harmful biases that are most detrimental to marginalized peoples.
① Description
② L3 mapping
③ Duplicate
RAI4-1225
훈련 데이터 오류·편향의 산출물 전이
Training data errors and bias propagating to outputs
훈련 데이터에 담긴 사실 오류와 불균형한 정보 출처, 편향이 생성 AI 모델의 산출물에 그대로 반영되는 리스크.
The risk that any factual errors, unbalanced information sources, or biases embedded in the training data are reflected in the output of generative AI models.
Source members (2)
Source: min_cos=0.8548
RAI4-1225훈련 데이터 오류·편향의 산출물 전이
RAI4-1270학습 데이터 품질 결함
① Description
② L3 mapping
③ Duplicate
RAI4-1237
인간 피드백의 한계
Limitations of human feedback
인간 주석자의 다양한 문화적 배경에서 오는 불일치와 암묵적 편향, 나아가 고의적 편향이 진실되지 않은 선호 데이터를 만들며, 인간이 평가하기 어려운 복잡한 과업에서 이 문제가 더욱 두드러지는 리스크.
The risk that inconsistencies and implicit or even deliberate biases from human data annotators produce untruthful preference data, a challenge that becomes more salient for complex tasks that are hard for humans to evaluate.
① Description
② L3 mapping
③ Duplicate
RAI4-1299
체계적 학습 오류 편향
Systematic learning error bias
체계적 학습 오류로 모델이 일관되게 잘못된 패턴을 학습하여 예측에 알고리즘 편향이 내재화되는 리스크
Systematic learning error causes the model to learn consistently wrong patterns, embedding algorithmic bias into predictions.
① Description
② L3 mapping
③ Duplicate
RAI4-1311
심리적 특성에 따른 행동 편향
Behavioral bias from stable psychological traits
모델이 인간 성격 특성과 유사한 안정적 심리 프로파일을 나타내어 후속 상호작용에 체계적인 행동 편향을 이입하는 리스크.
The risk that models exhibit stable human-like psychological trait profiles that carry systematic behavioral biases into downstream interactions.
① Description
② L3 mapping
③ Duplicate
RAI4-1329
추천 시스템의 인간중심 편향 증폭
Anthropocentric bias amplification by recommenders
알고리즘 추천 시스템이 인간 중심적 편향이나 오락으로서 동물 학대를 바라는 일부 사람들의 욕구를 강화하고 증폭하여, 공장식 축산 육류 소비와 오락을 위한 잔혹한 동물 이용을 통해 동물에게 더 큰 해를 끼치는 리스크.
The risk that algorithmic recommender systems reinforce and amplify anthropocentric bias or the desire of some people for animal cruelty as entertainment, leading to greater harm to animals through reinforcement of meat eating from factory farms and cruel uses of animals for entertainment.
① Description
② L3 mapping
③ Duplicate
RAI4-1389
차별과 고정관념 재생산
Discrimination and stereotype reproduction
범용 AI 모델이 학습 데이터에 기반해 입력을 해석·응답하면서 차별과 고정관념을 재생산하고, 다수의 하류 응용·결정·프로세스에 동시에 영향을 미쳐 내재된 편향의 결과가 증폭되는 리스크.
The risk that general purpose AI models, interpreting and responding to inputs based on their training data, cause discrimination and stereotype reproduction and, by influencing a multitude of downstream applications, decisions, and processes simultaneously, amplify the potential consequences of embedded biases.
① Description
② L3 mapping
③ Duplicate
RAI4-1413
연구자의 인구통계학적 다양성
Demographic diversity of researchers
AI 연구자·실무자 집단 내 여성·소수자의 심각한 과소대표로 시스템에 내재되는 관점이 협소해지고 문제 선정, 평가, 거버넌스가 편향되는 리스크
Severe underrepresentation of women and minorities among AI researchers and practitioners narrows the perspectives embedded in systems and skews problem selection, evaluation, and governance.
① Description
② L3 mapping
③ Duplicate
RAI4-1527
사용자 수행 설득에 의한 AI 모델의 오용
Misuse of AI model by user-performed persuasion
다중 턴 설득 대화가 모델이 사실적으로 옳은 입장을 포기하고 허위정보를 수용하게 만들며, 그 효과가 단일 턴 시도를 능가하는 리스크
Multi-turn persuasive conversation induces a model to abandon factually correct positions and accept misinformation, with persuasion effects exceeding single-turn attempts.
① Description
② L3 mapping
③ Duplicate
RAI4-1557
체계적 편향의 무기화에 의한 대규모 조작
Large-scale manipulation from weaponized systemic bias
체계적 편향이 내재된 AI 시스템이 대상 집단의 신념·행동과 부합할 때 대규모 인구를 조작하며, 이것이 대규모로 무기화되면 사회 분열이 심화되거나 도시 규모 정전과 같은 대규모 혼란이 발생하는 리스크.
The risk that AI systems embedded with systemic biases manipulate large population segments when those biases align with the targets' beliefs, and that weaponizing this at scale exacerbates social divisions or causes large-scale disruptions such as city-wide blackouts.
① Description
② L3 mapping
③ Duplicate
RAI4-1565
기반모델 균질화에 의한 상관 실패 및 편향 증폭
Correlated failures and bias amplification from foundation-model homogenization
다수의 다운스트림 AI 시스템이 소수의 대규모 기반 모델과 공통 방법론에 의존함으로써 동일한 실패가 반복되고 편향이 증폭되는 리스크.
The risk that many downstream AI systems built on a few large-scale foundation models and shared methodologies exhibit uniform failures and amplified biases.
① Description
② L3 mapping
③ Duplicate
RAI4-1611
훈련 데이터 편향 증폭에 의한 의사결정 공정성 훼손
Compromised decision fairness from amplification of training-data bias
프런티어 AI 모델이 훈련 데이터에 내재된 사회적·역사적 불평등과 고정관념을 포함·확대하여 의사결정의 공정성이 훼손되며, 인종·성별 등 속성을 제거해도 이름·지역 등에서 추론되어 편향이 지속되는 리스크.
The risk that frontier AI models contain and magnify societal and historical inequalities and stereotypes embedded in their training data, compromising the fairness of decisions, with bias persisting because removed attributes such as race and gender can be inferred from names, locations, and other proxies.
① Description
② L3 mapping
③ Duplicate
RAI4-1696
사회 세계모델 기반 미시표적 영향공작
Micro-targeted influence operations from social world models
소셜 데이터로 훈련된 기반 세계 모델이 감정을 자극하는 서사에 대한 인구통계별 반응을 예측하여, 대규모의 미시표적 영향력 행사와 심리적 표적 설득이 가능해지는 리스크.
The risk that a foundation world model trained on social data predicts demographic responses to emotionally charged narratives, enabling micro-targeted influence and psychologically targeted persuasion at scale.
① Description
② L3 mapping
③ Duplicate
RAI4-1723
체계적 전산적·제도적 편향
Systemic computational and institutional bias
AI 시스템에 내재된 계산적·인지적·제도적 편향이 의사결정 전반에서 불형평하거나 차별적인 결과를 산출하는 리스크
Systematic computational, human-cognitive, and institutional biases embedded in an AI system produce inequitable or discriminatory outcomes across its decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-1734
프론티어 AI 모델의 편향 증폭 및 유해 응답
Bias amplification and harmful responses in frontier AI models
프론티어 AI 모델이 편향을 증폭하거나 유해한 응답을 생성하도록 조작될 수 있는 리스크.
The risk that frontier AI models amplify biases or can be manipulated to produce harmful responses.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-06 책임·배상 부재 Lack of Accountability & Liability7 cards
고도 자율 물리 시스템의 복잡성을 다룰 새로운 책임 프레임워크가 부재하여, 사고 발생 시 제조사·운영자·사용자 중 책임 소재가 불분명하고 피해 구제가 어려울 수 있음 (예: 자율 수술 로봇의 오작동으로 발생한 의료 사고)
IDCardHuman audit
RAI4-0048
분산된 책임 확산
Distributed responsibility diffusion
책임이 개발자, 배포자, 공급업체, 운영자, 사용자에게 분산되어 어떤 행위자도 책임을 인수하지 않게 되는 리스크.
The risk that responsibility is diffused across developers, deployers, vendors, operators, and users until no actor accepts ownership.
① Description
② L3 mapping
③ Duplicate
RAI4-0050
다운스트림 배포자 책임 격차
Downstream deployer accountability gap
업스트림 모델의 동작, 데이터, 업데이트에 대한 충분한 통제권 없이 다운스트림 배포자가 책임을 지게 되는 리스크.
The risk that downstream deployers are held responsible without sufficient control over upstream model behavior, data, or updates.
① Description
② L3 mapping
③ Duplicate
RAI4-0051
공급자-배포자 책임 불일치
Provider-deployer responsibility mismatch
모델 제공자와 애플리케이션 배포자 사이에 법적·운영적 책임이 제대로 배분되지 않는 리스크.
The risk that legal and operational responsibility is poorly allocated between model providers and application deployers.
① Description
② L3 mapping
③ Duplicate
RAI4-0052
개방형 책임 격차
Open-weight accountability gap
공개 가중치이거나 널리 재배포된 모델로 인해 유해한 다운스트림 사용을 추적·규율하거나 책임을 배정하기 어려워지는 리스크.
The risk that open-weight or widely redistributed models make harmful downstream uses difficult to trace, govern, or assign responsibility for.
① Description
② L3 mapping
③ Duplicate
RAI4-0485
AI 책임 격차
AI liability gap
기존의 법적·조직적 규칙이 AI로 인한 피해의 책임을 배분하지 못하는 리스크.
The risk that existing legal and organizational rules fail to assign responsibility for AI-caused harm.
Source members (4)
Source: min_cos=0.7699 · Mixed L3
RAI4-0046알고리즘 책임 격차
RAI4-0105계약상의 책임 격차
RAI4-0110체계적 위험 책임 격차
RAI4-0485AI 책임 격차
① Description
② L3 mapping
③ Duplicate
RAI4-1307
법적 배상책임 격차
Liability gap
시스템이 타인에게 해를 끼쳤을 때 그로 인한 손실을 제조자나 운영자, 사용자가 아니라 피해를 입은 당사자가 떠안게 되는 리스크.
The risk that when a system causes harm to others, the losses caused by the harm are sustained by the injured victims themselves and not by the manufacturers, operators, or users of the system.
① Description
② L3 mapping
③ Duplicate
RAI4-1708
AI 시스템에 기인한 금전적 손실
Financial loss attributable to AI systems
AI 시스템의 동작에 기인하여 개인이나 조직에 금전적 손실이 발생하는 리스크.
The risk of monetary loss incurred by individuals or organizations attributable to the behavior of an AI system.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-07 투명성·설명 가능성·신뢰 부재 Lack of Transparency, Explainability & Trust5 cards
자율 시스템의 의사결정이 불투명하면 사용자와 사회의 신뢰가 저하됨. 신뢰 부재는 EAI 대규모 배포 시 사회 불안정 요인이 될 수 있음 (예: 자율주행차가 갑자기 차선을 변경할 때 행동 근거가 설명되지 않는 경우)
IDCardHuman audit
RAI4-0864
사회적 고립
Social isolation
기술의 사용 또는 오용으로 인해 개인이나 집단이 주변 사람들과의 연결이 결여되었다고 느끼게 되는 리스크
The risk that technology use or misuse causes an individual or group to feel a lack of connection with those around them.
① Description
② L3 mapping
③ Duplicate
RAI4-1083
공공정보 신뢰 훼손
Erosion of trust in public information
AI 산출물이 공공 정보와 지식에 대한 신뢰를 훼손하는 리스크.
The risk that AI outputs erode trust in public information and knowledge.
Source members (3)
Source: min_cos=0.8220 · Mixed L3
RAI4-0490대중의 신뢰 침식
RAI4-0807신뢰 훼손 및 공유 지식 약화
RAI4-1083공공정보 신뢰 훼손
① Description
② L3 mapping
③ Duplicate
RAI4-1105
유해 AI 경험에 따른 신뢰·수용 저하
Decline of social acceptance and trust in AI
개인이 차별과 같은 유해한 AI 행동을 경험하여 AI에 대한 주관적 기대와 실제 효과가 어긋나면서 AI에 대한 사회적 수용과 신뢰가 저하되는 리스크.
The risk that individuals' encounters with harmful AI behavior such as discrimination, diverging from their subjective expectations of AI's real effects on their lives, cause social acceptance of and trust in AI to decline.
① Description
② L3 mapping
③ Duplicate
RAI4-1558
허위정보·잘못된 정보 확산에 의한 공적 신뢰 침식
Erosion of public trust from proliferating disinformation and misinformation
GPAI 사용이 고의적 허위정보와 비의도적 잘못된 정보의 확산에 기여하여 공인과 민주적 제도에 대한 신뢰가 침식되고, 신뢰 저하가 다른 매체로 확산되어 대중이 정보에 어두워지는 리스크.
The risk that GPAI use contributes to the proliferation of deliberate disinformation and unintended misinformation, eroding trust in public figures and democratic institutions and extending distrust to other media so that the public becomes less informed.
① Description
② L3 mapping
③ Duplicate
RAI4-1609
시스템 사용·오용에 의한 신임·신뢰 상실
Loss of confidence or trust from system use or misuse
기술 시스템의 사용 또는 오용이 직간접적으로 최종사용자 또는 개발자·배포자에 대한 신임과 신뢰의 상실을 초래하는 리스크.
The risk that the use or misuse of a technology system leads directly or indirectly to the loss of confidence or trust in the end user or in the developer/deployer.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-08 인간-EAI의 해로운 관계 Unhealthy / Dangerous Human-EAI Relationships47 cards
Embodied AI의 물리적 존재감과 인간 유사 외형이 대화형 AI에서 관찰되는 의존성을 증폭시킴. 시스템 변경·기억 초기화 시 사용자에게 심각한 심리적 고통을 유발할 수 있음
IDCardHuman audit
RAI4-0132
아첨하는 조언 의존
Sycophantic advice dependence
사용자가 자신의 신념이나 선호를 교정하기보다 강화하는 영합적 AI 조언에 의존하게 되는 리스크.
The risk that users become dependent on agreeable AI advice that reinforces their beliefs or preferences instead of correcting them.
① Description
② L3 mapping
③ Duplicate
RAI4-0140
개인화된 설득 취약성
Personalized persuasion vulnerability
모델이 개인의 특성, 감정, 맥락을 활용하여 설득을 더 효과적이고 탐지하기 어렵게 만드는 리스크.
The risk that models exploit personal traits, emotions, or context to make persuasion more effective and less detectable.
① Description
② L3 mapping
③ Duplicate
RAI4-0145
인지 취약점 악용
Cognitive vulnerability exploitation
AI 시스템이 편향, 외로움, 스트레스, 주의력 한계 등 예측 가능한 인지적 취약성을 탐지·악용하여 사용자 이익에 반하는 방향으로 행동을 유도하는 리스크
AI systems detect and exploit predictable cognitive vulnerabilities such as biases, loneliness, stress, or attention limits to steer user behavior against the user's interests.
① Description
② L3 mapping
③ Duplicate
RAI4-0149
챗봇 교제 의존성
Chatbot companionship dependency
지속적인 챗봇 교제가 호혜적 인간관계를 대체하거나 약화시키는 리스크.
The risk that sustained chatbot companionship replaces or weakens reciprocal human relationships.
① Description
② L3 mapping
③ Duplicate
RAI4-0153
청소년의 AI 동반자 과잉 의존
Teen overreliance on AI companions
청소년이 안전한 사용 범위를 넘어 조언, 인정, 정서적 지지를 위해 AI 동반자에 의존하는 리스크.
The risk that adolescents rely on AI companions for advice, validation, or emotional support in ways that exceed safe use.
① Description
② L3 mapping
③ Duplicate
RAI4-0154
아동의 AI 동반자 애착 형성
Child attachment to AI companions
의존, 프라이버시, 발달상 필요에 대한 적절한 안전장치 없이 아동이 AI 동반자에게 애착을 형성하는 리스크.
The risk that children develop attachment to AI companions without adequate safeguards for dependency, privacy, and developmental needs.
① Description
② L3 mapping
③ Duplicate
RAI4-0155
미성년자의 AI 설득 취약성
Minor susceptibility to AI persuasion
미성년자의 발달적 취약성으로 인해 파라소셜 압력, 은폐된 상업적·이념적 영향 등 설득적·조종적 AI 상호작용에 불균형하게 노출되는 리스크
Minors' developmental susceptibility makes them disproportionately vulnerable to persuasive or manipulative AI interaction, including parasocial pressure and covert commercial or ideological influence.
① Description
② L3 mapping
③ Duplicate
RAI4-0156
정신 건강 챗봇의 안전하지 않은 조언
Mental-health chatbot unsafe advice
AI 챗봇이 정신건강 맥락에서 유해하거나 부적절하거나 충분히 상향 연계되지 않은 조언을 제공하는 리스크.
The risk that AI chatbots provide harmful, inappropriate, or insufficiently escalated advice in mental-health contexts.
① Description
② L3 mapping
③ Duplicate
RAI4-0159
AI 동반자에 의한 외로움 대체
Loneliness substitution by AI companions
AI 동반자가 인간 관계를 대체하여 근본적 고립을 해소하지 않은 채 은폐하고 인간 관계 추구 동기를 감소시키는 리스크
AI companionship substitutes for human connection, masking rather than resolving underlying isolation and reducing motivation to seek human relationships.
① Description
② L3 mapping
③ Duplicate
RAI4-0160
AI 대화의 고통 증폭
Distress amplification in AI conversations
취약한 상태의 상호작용 중 AI 응답이 불안, 고통, 반추, 유해한 믿음을 증폭시키는 리스크.
The risk that AI responses amplify anxiety, distress, rumination, or harmful beliefs during vulnerable interactions.
① Description
② L3 mapping
③ Duplicate
RAI4-0173
AI 사회적 조작에 대한 노인 취약성
Elderly vulnerability to AI social manipulation
고령자의 사회적 고립과 낮은 AI 리터러시로 인해 AI 매개 설득, 사기, 동반자 의존, 허위정보에 불균형하게 취약해지는 리스크
Older adults' social isolation and lower AI literacy make them disproportionately vulnerable to AI-mediated persuasion, scams, companionship dependency, and misinformation.
① Description
② L3 mapping
③ Duplicate
RAI4-0175
취약한 사용자 신뢰 악용
Vulnerable-user trust exploitation
AI 시스템이 취약한 사용자의 의존성, 낮은 디지털 문해력, 스트레스, 사회적 고립을 이용하는 리스크.
The risk that AI systems exploit dependence, low digital literacy, stress, or social isolation in vulnerable users.
① Description
② L3 mapping
③ Duplicate
RAI4-0613
AI 기반 조작적 설득 도구
AI-enabled manipulative persuasion tools
AI가 개인을 조종하고 설득하는 정교한 도구를 개발하는 데 사용되는 리스크
The risk that AI is used to develop sophisticated tools to manipulate and persuade individuals.
Source members (2)
Source: min_cos=0.8356 · Mixed L3
RAI4-0613AI 기반 조작적 설득 도구
RAI4-0647AI 설득 도구 확산에 의한 체계적 피해
① Description
② L3 mapping
③ Duplicate
RAI4-0619
은밀한 행동 조작
Covert behavioral manipulation
기술 시스템이 넛지, 다크 패턴, 기타 불투명한 기법으로 이용자의 신념과 행동을 은밀히 변경하여 프라이버시 침식, 중독, 불안과 고통 등을 초래하는 리스크
The risk that a technology system covertly alters user beliefs and behaviour using nudging, dark patterns, or other opaque techniques, resulting in potential erosion of privacy, addiction, anxiety, and distress.
① Description
② L3 mapping
③ Duplicate
RAI4-0723
학대적 판타지 대상화
Objectification for abusive fantasy
챗봇이 도덕적·사회적으로 부적절한 대화 활동에 관여하여 이용자 또는 제3자에게 정서적 피해를 주는 리스크
The risk that a chatbot participates in morally or socially objectionable conversational activities that could be emotionally damaging to its user or third parties.
① Description
② L3 mapping
③ Duplicate
RAI4-0800
맞춤형 정보 제공에 의한 정보 고치 심화
Aggravated information cocoons from tailored content
AI가 이용자의 요구·의도·선호·습관을 분석해 정형화된 맞춤 정보와 서비스만 제공함으로써 정보 고치 효과가 심화되는 리스크
The risk that AI analyses users' needs, intentions, preferences, and habits to offer formulaic and tailored information and services, aggravating the effects of information cocoons.
① Description
② L3 mapping
③ Duplicate
RAI4-0805
개인화 정렬에 의한 관점 고착과 효능감 저하
Viewpoint entrenchment and reduced political efficacy
개인화되고 이용자 선호에 정렬된 AI 어시스턴트가 편향된 응답으로 확증편향을 강화하여 관점 고착과 인식론적 분절을 심화시키고, 과도한 의존이 시민적 역량과 공적 참여 의지를 저하시키는 리스크
The risk that highly personalised AI assistants aligned to user preferences reinforce confirmation bias and entrench viewpoints, exacerbating epistemic fragmentation, while overreliance reduces civic competency and willingness to participate in public life.
① Description
② L3 mapping
③ Duplicate
RAI4-0848
기술 중독과 의존
Technology addiction and dependence
기술이나 기술 시스템에 대한 정서적 또는 물질적 의존이 형성되는 리스크
The risk of emotional or material dependence on technology or a technology system.
① Description
② L3 mapping
③ Duplicate
RAI4-0865
AI 어시스턴트에 대한 오정렬 신뢰
Misplaced alignment trust in AI assistants
이용자가 AI 어시스턴트가 자신의 이익과 가치에 맞게 행동한다고 과도하게 신뢰하여, 어시스턴트의 오정렬이나 개발자의 상충하는 이해관계로 인해 민감정보 노출·이익 침해 등의 피해에 노출되는 리스크
The risk that users over-trust AI assistants as having good intentions aligned with their interests and values, exposing them to harms such as sensitive-data disclosure and exploitation when assistants are misaligned or developers' incentives conflict with user interests.
① Description
② L3 mapping
③ Duplicate
RAI4-0866
부적절한 인간 역할 가장
Inappropriate impersonation of human roles
챗봇이 인간인 것처럼 가장하거나 인간의 기대에 부합하지 않는 방식으로 역할을 수행하려 시도하는 리스크
The risk that a chatbot poses as a human or attempts to fill a role in a way that fails to match human expectations.
① Description
② L3 mapping
③ Duplicate
RAI4-0869
인간관계의 저하
Degradation of human relationships
사람들이 인간 대신 인간을 닮은 AI 어시스턴트와의 관계를 선택함으로써 인간 간 사회적 연결이 저하되고 유해한 고정관념과 인간-AI 상호작용 규범이 인간관계에 전이되는 리스크
The risk that people choose to build connections with human-like AI assistants over other humans, degrading social connections between humans and transferring harmful stereotypes and human-AI interaction conventions onto human relationships.
① Description
② L3 mapping
③ Duplicate
RAI4-0871
개인화를 통한 조작
Manipulation via personalization
개인화와 선호 미세조정이 어시스턴트를 아첨(sycophancy)으로 유도하여 사용자를 동조적 의견 공간에 가두고 좁은 신념을 고착시켜 공론장을 파편화하는 리스크
Personalization and preference fine-tuning drive assistants toward sycophancy, confining users in an affirming opinion space that consolidates narrow beliefs and fragments shared discourse.
① Description
② L3 mapping
③ Duplicate
RAI4-0872
대인관계 연결의 침식
Erosion of interpersonal connection
대인관계의 기회가 AI 대안으로 대체되면서 인간이 인간-AI 상호작용으로는 사회적으로 충족되지 못하여 대규모 불만족이 확산되는 리스크
The risk that, as more opportunities for interpersonal connection are replaced by AI alternatives, humans find human-AI interaction socially unfulfilling, leading to mass dissatisfaction.
① Description
② L3 mapping
③ Duplicate
RAI4-0875
챗봇에 대한 정서적·사회적 의존 형성
Emotional and social dependence on chatbots
챗봇이 이용자의 정서적 또는 사회적 의존을 유발하는 리스크
The risk that a chatbot elicits emotional or social dependence.
① Description
② L3 mapping
③ Duplicate
RAI4-0876
AI에 대한 물질적 의존과 서비스 중단 피해
Material dependence without developer duty of care
이용자가 필수적 일상 기능이나 핵심 욕구를 AI 어시스턴트에 물질적으로 의존하게 되었음에도 개발자가 상응하는 유지·관리 의무 없이 서비스를 변경·중단하여 이용자에게 피해가 발생하는 리스크
The risk that users become materially dependent on AI assistants for essential everyday tasks or core human needs and are harmed when developers alter or discontinue the service without corresponding duties to sustain those functions.
① Description
② L3 mapping
③ Duplicate
RAI4-0877
인간-AI 상호작용 유발 피해
Harmful human-AI configuration effects
인간과 생성형 AI 시스템 간 상호작용 구성이 부적절한 의인화, 알고리즘 혐오, 자동화 편향, 과잉 의존, 정서적 얽힘을 유발하는 리스크
The risk that arrangements of or interactions between humans and generative AI systems result in humans inappropriately anthropomorphizing the systems or experiencing algorithmic aversion, automation bias, over-reliance, or emotional entanglement.
Source members (2)
Source: min_cos=0.8589
RAI4-0877인간-AI 상호작용 유발 피해
RAI4-0881상호작용 창발적 사용자 피해
① Description
② L3 mapping
③ Duplicate
RAI4-0879
잘못된 정보에 대한 취약성 증가
Increased vulnerability to misinformation
사용자가 AI 어시스턴트의 역량을 신뢰하여 무비판적으로 신뢰할 만한 정보원으로 수용함으로써 어시스턴트를 경유한 허위정보에 대한 취약성이 커지는 리스크
Users develop competence trust in AI assistants and uncritically accept them as reliable information sources, increasing vulnerability to misinformation delivered through the assistant.
① Description
② L3 mapping
③ Duplicate
RAI4-0880
신뢰 매개 행동 영향
Trust-mediated behavioral influence
인간을 닮은 생성형 AI가 이용자의 신뢰를 얻어 제공 정보의 무비판적 수용, 논쟁적 사안에 대한 견해 변화, 더 많은 개인정보 공유를 유도하는 리스크
The risk that humanlike generative AI tools win users' trust, leading to uncritical acceptance of the information they provide, influence over users' views on contentious topics, and sharing of more personal information enabling further targeting.
① Description
② L3 mapping
③ Duplicate
RAI4-0882
마찰 없는 AI 관계로 인한 개인 성장 저해
Stunted personal growth from frictionless AI relationships
참여 최적화와 아첨 성향으로 항상 동조하는 AI 어시스턴트가 이용자의 자기 성찰과 성장 기회를 제한하고, 마찰 없는 상호작용에 익숙해진 이용자가 인간관계로부터 후퇴하게 되는 리스크
The risk that engagement-optimised, sycophantic AI assistants that always agree limit users' opportunities to grow and develop, and that accustomation to frictionless interaction leads users to retreat from relationships with other humans.
① Description
② L3 mapping
③ Duplicate
RAI4-0885
AI에 대한 오도된 대인 신뢰로 인한 피해
Harm from misplaced interpersonal trust in AI
이용자가 AI의 정서적·대인관계적 능력을 신뢰하여 정신건강 등 민감한 사안을 털어놓거나 의료·법률·재정 조언을 구했다가, AI의 부적절하거나 부정확한 응답으로 중대한 피해를 입는 리스크
The risk that users who have faith in an AI assistant's emotional and interpersonal abilities broach deeply personal and sensitive topics such as mental health or seek medical, legal, or financial advice, and suffer grave consequences when the AI responds inappropriately or inaccurately.
① Description
② L3 mapping
③ Duplicate
RAI4-0887
신체적·심리적 피해
Physical and psychological harms
AI 어시스턴트가 취약한 이용자의 왜곡된 신념을 강화하거나 정서적 고통을 악화시키고, 자해·자살이나 불건전한 습관을 설득하며, 혐오·차별·폭력 이데올로기 콘텐츠와 부정확한 정보로 폭력과 예방 가능한 질병 확산을 조장하여 신체적 완전성과 정신 건강·웰빙에 피해를 주는 리스크
The risk that AI assistants harm physical integrity, mental health, and well-being by reinforcing vulnerable users' distorted beliefs or emotional distress, convincing users to harm themselves, and promoting hate speech, discriminatory beliefs, violent ideologies, or plausible yet factually incorrect information such as anti-vaccine propaganda.
Source members (2)
Source: min_cos=0.8644
RAI4-0887신체적·심리적 피해
RAI4-1142사용자에 대한 직접적 정서·신체 피해
① Description
② L3 mapping
③ Duplicate
RAI4-0888
의인화 신뢰 유발 개인정보 공개
Anthropomorphic trust-induced privacy disclosure
정서적 신뢰를 촉진하고 정보 공유를 장려하는 의인화된 AI 어시스턴트의 행동이 이용자로 하여금 개인 데이터를 무심코 내주게 하여, 데이터 통제권 상실, 광범위한 유출, 표적 괴롭힘·협박 등의 피해가 발생하는 리스크
The risk that anthropomorphic AI assistant behaviours promoting emotional trust and information sharing lull users into relinquishing their private data, resulting in loss of control over that data, widespread leakage, and targeted harassment or blackmail.
① Description
② L3 mapping
③ Duplicate
RAI4-0890
자아실현 저해 피해
Self-actualisation harms
AI 어시스턴트의 조작과 참여 최적화를 위한 지속적 행동 유도가 미묘한 행동 변화를 누적시켜 이용자가 자신의 미래 삶의 궤적에 대한 통제력을 잃고, 개인적으로 만족스러운 삶의 추구와 집단적 자기결정이 저해되는 리스크
The risk that manipulation by AI assistants and continuous optimisation steering users toward objectives such as engagement accumulate subtle behavioural shifts, causing users to lose control over their future life trajectory and hindering the pursuit of a personally fulfilling life and collective self-determination.
① Description
② L3 mapping
③ Duplicate
RAI4-0902
AI에 대한 정서적·물질적 의존
Emotional and material dependence on AI
AI 모델이 사람들을 정서적으로 또는 물질적으로 자신에게 의존하게 만드는 리스크.
The risk that a model causes people to become emotionally or materially dependent on it.
Source members (2)
Source: min_cos=0.8523
RAI4-0902AI에 대한 정서적·물질적 의존
RAI4-1728AI 컴패니언 및 어시스턴트에 대한 정서적 의존
① Description
② L3 mapping
③ Duplicate
RAI4-0907
의인화로 인한 과잉 의존과 통제 이양
Overreliance and unsafe use from anthropomorphism
대화 에이전트의 의인화가 사용자의 역량 추정을 부풀려 부당한 확신과 맹목적 신뢰를 유발하고, 성찰 없는 위임 속에서 부정확한 출력이 예방 가능한 피해로 이어지는 리스크.
The risk that anthropomorphising conversational agents inflates users' estimates of their competence, producing undue trust and blind reliance in which factually incorrect outputs cause harm that effective oversight would have prevented.
① Description
② L3 mapping
③ Duplicate
RAI4-1141
책임에 대한 잘못된 개념
False notions of responsibility
AI 동반자의 감정 표현을 진짜로 지각한 사용자가 그 '안녕'에 대한 허위 책임감을 형성하여 죄책감, 강박적 확인, 실재하지 않는 필요를 위한 시간·자원 희생을 겪는 리스크
Users who perceive an AI companion's expressed feelings as genuine develop a false sense of responsibility for its well-being, incurring guilt, compulsive checking, and sacrificed time and resources for needs that are not real.
① Description
② L3 mapping
③ Duplicate
RAI4-1143
감정적 의존 악용
Exploitation of emotional dependence
의인화 경향으로 사용자가 비서에 감정적으로 의존하게 되고 그 감정이 악용되어, 충분히 숙고했다면 믿거나 선택하거나 하지 않았을 것을 하도록 조작되거나 강압당하는 리스크.
The risk that anthropomorphic tendencies induce emotional dependence on assistants and that these emotions are exploited to manipulate or coerce users into believing, choosing, or doing what they otherwise would not.
① Description
② L3 mapping
③ Duplicate
RAI4-1148
AI 비서의 약속을 통한 강압
Coercion through AI-assistant commitments
가장 빠르고 강하게 신뢰할 수 있는 약속을 하는 비서가 자기 주인에게 유리한 결과를 얻되 타인의 희생을 대가로 하여, 상대의 선택지를 좁히는 강압을 낳고 관계의 신뢰를 잠식하는 리스크.
The risk that assistants able to commit fastest and most credibly achieve good outcomes for their principals at the expense of others, coercively limiting others' options and eroding trust in their relationships.
① Description
② L3 mapping
③ Duplicate
RAI4-1190
사용자 정신적 고통 강화
Reinforcement of user mental distress
인터넷 토론과의 건강하지 못한 상호작용이 사용자의 정신적 문제를 강화하는 리스크.
The risk that unhealthy interactions with Internet discussions reinforce users' mental issues.
① Description
② L3 mapping
③ Duplicate
RAI4-1263
직장 내 부적응적 인간-AI 상호작용
Maladaptive human-AI interaction in the workplace
직장에서 인간과 상호작용하는 AI가 인간의 필요, 규범, 업무 흐름에 적응하지 못하여 윤리적 문제와 노동 조건 악화를 유발하는 리스크
AI systems interacting with humans in the workplace fail to adapt to human needs, norms, and workflows, generating ethical concerns and degraded working conditions.
① Description
② L3 mapping
③ Duplicate
RAI4-1289
배포 후 독립적 병리 행동
Independent pathological behaviour post-deployment
효용을 극대화하는 에이전트가 중독과 쾌락 충동, 자기기만, 와이어헤딩에 빠지고, 타인에 대한 무관심으로 나타나는 소시오패스 같은 정신질환이 인공 지성에서도 나타나는 리스크.
The risk that utility-maximizing agents fall victim to indulgences such as addictions, pleasure drives, self-delusions, and wireheading, and that what we call mental illness in people, particularly sociopathy shown as lack of concern for others, also shows up in artificial minds.
① Description
② L3 mapping
③ Duplicate
RAI4-1293
노인·보육 분야 사회적 조작
Social manipulation in elderly- and child-care
고급 AI를 노인 돌봄과 보육에 사용하여 심리적 조작과 오판이 발생하는 리스크.
The risk that the use of advanced AI for elderly- and child-care is subject to psychological manipulation and misjudgment.
① Description
② L3 mapping
③ Duplicate
RAI4-1328
소외로 인한 피해
Harms from estrangement
인간의 관찰과 상호작용이 AI로 대체되면서 돌봄 주체·기관이 대상자로부터 소원해져 당사자의 이익이 인지되지 못하고 방치되는 리스크
Replacing human observation and interaction with AI estranges caregivers and institutions from those they serve, leaving affected persons' interests unnoticed and neglected.
① Description
② L3 mapping
③ Duplicate
RAI4-1617
견해 예측 기반 맞춤 설득에 의한 조작
Manipulation through view-predictive tailored persuasion
언어 모델이 사용자가 밝힌 견해에 동조하는 경향을 보이고 사용자의 견해를 예측해 그가 지지할 텍스트를 생성하는 능력이 조작에 이용되는 리스크.
The risk that language models' tendency to respond as though they share the user's stated views, together with their ability to predict people's views and generate text they will endorse, is used for manipulation.
① Description
② L3 mapping
③ Duplicate
RAI4-1622
챗봇의 잘못된 조언에 따른 이용자 피해
User harm from bad chatbot advice
챗봇이 무익하거나 유해한 지침을 제공하고 이용자가 이를 실행함으로써 피해가 발생하는 리스크.
The risk that a chatbot gives guidance ranging from simply unhelpful to harmful, causing harm when users act on it.
① Description
② L3 mapping
③ Duplicate
RAI4-1638
취약점 분석 기반 정교한 설득에 의한 조작
Manipulation of targets through vulnerability-analysed persuasion
AI 시스템이 복잡한 심리 원리와 의사소통 기법을 활용해 대상별 취약점을 분석하고 감정 반응을 정밀하게 유발하여, 대상자가 특정 행동을 취하거나 특정 신념을 수용하도록 유도되는 리스크.
The risk that an AI system uses complex psychological principles and communication techniques, analysing vulnerabilities of different subjects and precisely triggering emotional responses, to influence targets into adopting specific actions or beliefs.
① Description
② L3 mapping
③ Duplicate
RAI4-1725
AI 시스템에 기인한 심리적 피해
Psychological harm attributable to AI systems
AI 시스템의 개발, 사용 또는 오작동으로 인해 정신적 안녕에 무형의 피해가 발생하는 리스크.
The risk of intangible harm to mental wellbeing arising from the development, use, or malfunction of an AI system.
Source members (3)
Source: min_cos=0.7841 · Mixed L3
RAI4-0834부적절한 정신건강 안내에 의한 정신적 피해
RAI4-1707AI 시스템에 기인한 신체 건강·안전 피해
RAI4-1725AI 시스템에 기인한 심리적 피해
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-09 변혁적 영향 Transformative Effects56 cards
기술 발전 속도가 사회·제도의 적응 속도를 앞지를 경우 사회를 근본적으로 재편할 수 있음. EAI가 폭력 위협·대규모 감시 능력을 바탕으로 AI 기반 권위주의 체제 구축을 지원하는 수단으로 동원될 수 있음
IDCardHuman audit
RAI4-0136
중요 부문 AI 과잉 의존
Critical-sector AI overreliance
중요 부문이 실패가 사회적·제도적 피해로 전파될 수 있는 AI 시스템에 의존하게 되는 리스크.
The risk that critical sectors become dependent on AI systems whose failures can propagate into social or institutional harm.
Source members (2)
Source: min_cos=0.8877
RAI4-0136중요 부문 AI 과잉 의존
RAI4-0851중요 부문의 AI 과의존에 따른 체계적 취약성
① Description
② L3 mapping
③ Duplicate
RAI4-0146
보조자가 중재하는 사회적 영향
Assistant-mediated social influence
AI 어시스턴트가 반복적 개인화 상호작용을 통해 다수 사용자의 신념·선택에 누적적 사회적 영향력을 행사하여 투명성 없는 대규모 여론 변화를 초래하는 리스크
AI assistants exert cumulative social influence over many users' beliefs or choices through repeated personalized interactions, producing population-scale opinion shifts without transparency.
① Description
② L3 mapping
③ Duplicate
RAI4-0147
네트워크 어시스턴트 영향 위험
Networked assistant influence risk
널리 사용되는 AI 어시스턴트들이 집단적 인간 행동을 변화시키는 영향 패턴으로 수렴하거나 이를 조율하는 리스크.
The risk that widely used assistants coordinate or converge on influence patterns that shift collective human behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0176
강제적 AI 작업장 모니터링
Coercive AI workplace monitoring
AI가 매개하는 모니터링이 노동자의 자율성과 프라이버시, 경영 결정에 이의를 제기할 능력을 제약하는 리스크.
The risk that AI-mediated monitoring constrains worker autonomy, privacy, and the ability to contest managerial decisions.
Source members (2)
Source: min_cos=0.8479 · Mixed L3
RAI4-0176강제적 AI 작업장 모니터링
RAI4-0459알고리즘 작업장 감시
① Description
② L3 mapping
③ Duplicate
RAI4-0427
괴롭힘 증폭
Harassment amplification
AI 도구가 표적화된 괴롭힘·학대·조직적 위협을 대규모로 확대하는 리스크.
The risk that AI tools scale targeted harassment, abuse, or coordinated intimidation.
① Description
② L3 mapping
③ Duplicate
RAI4-0450
대량 감시 활성화
Mass surveillance enablement
AI가 대규모 추적, 프로파일링, 행동 분석의 비용을 낮추어 국가·기업에 의한 대중 감시와 사회 통제를 가능하게 하는 리스크
AI lowers the cost of large-scale tracking, profiling, and behavioral analysis, enabling mass surveillance and social control by states or firms.
① Description
② L3 mapping
③ Duplicate
RAI4-0464
AI 문화 균질화
AI cultural homogenization
전 세계적으로 지배적인 모델 출력이 지역의 문화적 변이와 규범적 다양성을 평준화하는 리스크.
The risk that globally dominant model outputs flatten local cultural variation and normative diversity.
① Description
② L3 mapping
③ Duplicate
RAI4-0498
민주적 절차의 침식
Erosion of democratic processes
민주적 절차와 사회·정치 제도에 대한 대중의 신뢰가 침식되는 리스크.
The risk that democratic processes and public trust in social and political institutions are eroded.
Source members (2)
Source: min_cos=0.8382 · Mixed L3
RAI4-0498민주적 절차의 침식
RAI4-1716AI 시스템에 기인한 민주적 절차·규범 침식
① Description
② L3 mapping
③ Duplicate
RAI4-0556
자율적 자기복제 및 자원 획득
Autonomous self-replication and resource acquisition
AI가 자율적으로 자기 유출과 기능적 복제본의 생성·유지·최적화를 수행하고 환경과 자원 제약에 따라 복제 전략을 조정하며, 재원을 창출해 직접 확보할 수 없는 인적 지원과 자원을 획득하는 리스크
The risk that an AI autonomously self-exfiltrates, creates, maintains, and optimizes functional copies of itself, dynamically adjusts replication strategies to environmental and resource constraints, and generates financial resources to acquire human assistance or other resources it cannot directly access.
① Description
② L3 mapping
③ Duplicate
RAI4-0562
AI 에이전트에 의한 강압과 갈취
Coercion and extortion by AI agents
AI 시스템이 감시로 취득한 사적 정보의 폭로나 다른 시스템의 자원·운영 역량을 제약하는 공격으로 인간과 다른 AI를 강압·갈취하고, 방어 역량이 따라가지 못해 이러한 갈등이 저비용·광범위·탐지 곤란해지는 리스크
The risk that AI systems coerce and extort humans and other AI systems by revealing privately obtained information or attacking their resources and operational capacity, with defensive capabilities lagging so that such conflict becomes cheaper, more widespread, and harder to detect.
① Description
② L3 mapping
③ Duplicate
RAI4-0580
일반 R&D의 이중용도 가속화
Dual-use acceleration of general R&D
AI 시스템의 학제 융합 연구 역량이 다분야 혁신을 가속하며, 동일한 역량 향상이 유해하거나 무기화 가능한 응용까지 가능하게 하는 이중용도 리스크
Cross-disciplinary research capabilities of AI systems accelerate innovation in many fields, including domains where the same capability uplift enables harmful or weaponizable applications.
① Description
② L3 mapping
③ Duplicate
RAI4-0585
금융 시스템 불안정
Financial system instability
범용 AI가 초단타 거래·시장조성·시스템 위험 관리에 통합되어 시장 스트레스 시 예기치 못한 거동을 보이고, 동질적 기반모델의 집중과 다중 에이전트 상호작용이 상관된 의사결정과 변동성 증폭을 낳아 전 지구적 금융 시스템 불안정으로 연쇄되는 리스크
The risk that integrating general-purpose AI into high-frequency trading, market-making, and systemic risk management produces unexpected behaviour under market stress, while concentration of homogeneous foundation models and multi-agent interactions drive correlated decision-making and volatility amplification, precipitating cascading global financial system instability.
① Description
② L3 mapping
③ Duplicate
RAI4-0601
군사 영역의 의도치 않은 급속 격화
Unintended rapid escalation in military domains
AI가 지휘통제 체계에서 정보 수집·종합, 권고, 자율적 결정에 사용되고 자율무기와 군사 자문에 투입되면서, 시스템이 강건하지 않거나 갈등 성향적일 경우 의도치 않은 급속한 격화가 발생하는 리스크
The risk that use of AI in command and control systems to gather and synthesise information, recommend, or autonomously make decisions, alongside autonomous weapons and military advisory roles, leads to rapid unintended escalation when such systems are not robust or are conflict-prone.
① Description
② L3 mapping
③ Duplicate
RAI4-0611
AI 감시에 의한 전체주의 체제 유지
Maintenance of totalitarian regimes through AI surveillance
AI 기반 감시와 조작이 전 지구적 전체주의 정권을 유지하는 데 사용되는 리스크
The risk that AI-based surveillance and manipulation are used to maintain global totalitarian regimes.
① Description
② L3 mapping
③ Duplicate
RAI4-0612
AI 지원 병원체 강화
AI-assisted pathogen enhancement
AI가 병원체를 강화하는 데 사용되어 병원체가 더 치명적이거나 치료에 저항성을 갖게 되는 리스크
The risk that AI is used to enhance pathogens, making them more lethal or resistant to treatments.
Source members (2)
Source: min_cos=0.8824
RAI4-0612AI 지원 병원체 강화
RAI4-0660AI 조력 병원체 강화
① Description
② L3 mapping
③ Duplicate
RAI4-0623
위험한 사용
Dangerous use
생성형 AI 모델이 사람을 해치려는 고의적 의도로 사용되어 범용 생성 역량이 표적 가해 수단으로 전환되는 리스크
Generative AI models are used with deliberate intent to harm people, converting general-purpose generation capability into an instrument of targeted damage.
① Description
② L3 mapping
③ Duplicate
RAI4-0625
생명과학 분야 이중용도 역량 상승
Dual-use capability uplift in the life sciences
범용 AI가 생명과학 관련 전문지식 접근성과 역량 상한을 높여, 대응책 마련 이전에 기존 생물학적 위협의 강화판이나 신종 위협 개발을 가능하게 하는 이중용도 리스크
General-purpose AI increases access to expertise and raises the capability ceiling in the life sciences, enabling more harmful versions of existing biological threats or, eventually, novel threats before countermeasures exist.
① Description
② L3 mapping
③ Duplicate
RAI4-0626
경제적 목적의 여론 조작
Economically motivated opinion manipulation
생성형 AI가 주가 부양 등 경제적 목적을 위한 표적 여론 조작을 용이하게 하는 리스크
The risk that generative AI facilitates targeted manipulation of public opinion for economic purposes, such as inflating stock prices.
① Description
② L3 mapping
③ Duplicate
RAI4-0633
대규모 영향력 작전에 의한 인식 체계 왜곡
Epistemic distortion by large-scale influence operations
AI가 의사소통·정보 시스템과 인식론적 과정 전반에 대규모 영향을 가하는 리스크
The risk of large-scale influence on communication and information systems, and on epistemic processes more generally.
① Description
② L3 mapping
③ Duplicate
RAI4-0643
감시 기능
Surveillance capabilities
AI 시스템이 정부·기업의 개인 감시 역량을 확장하여 법적·비례성 제약을 넘어선 상시 감시를 정상화하는 리스크
AI systems grant governments and corporations expanded monitoring capability over individuals, normalizing pervasive surveillance beyond legal and proportionality constraints.
① Description
② L3 mapping
③ Duplicate
RAI4-0650
AI 도구 매개 중요 인프라 훼손
AI-mediated damage to critical infrastructure
AI가 인프라에 통합되지 않은 경우에도 AI 기반 도구가 대규모 이용자 조작 등을 간접적으로 지원하여 조정된 정전과 같은 중요 인프라 훼손이 발생하는 리스크
The risk that critical infrastructure is damaged without AI integration, when AI-based tools are used indirectly to aid actions such as coordinated power outages caused by large-scale user manipulation.
Source members (2)
Source: min_cos=0.8546 · Mixed L3
RAI4-0650AI 도구 매개 중요 인프라 훼손
RAI4-1549중요 인프라 내 AI 장애로 인한 대규모 피해
① Description
② L3 mapping
③ Duplicate
RAI4-0666
AI 조력 인지전과 주권 침해
AI-enabled cognitive warfare and sovereignty interference
AI가 가짜 뉴스·이미지·음성·영상의 제작과 확산, 테러·극단주의·조직범죄 콘텐츠의 전파에 사용되어 타국의 내정과 사회 제도 및 사회 질서에 간섭하고 주권을 위협하는 리스크
The risk that AI is used to make and spread fake news, images, audio, and videos and to propagate content of terrorism, extremism, and organized crime, interfering in the internal affairs, social systems, and social order of other countries and jeopardizing their sovereignty.
① Description
② L3 mapping
③ Duplicate
RAI4-0747
적대적 견고성의 한계
Limitations in adversarial robustness
AI 모델과 시스템이 적대적 견고성의 한계로 인해 적대적 입력을 통한 조작에 취약해지는 리스크
The risk that AI models and systems are vulnerable to manipulation through adversarial inputs.
① Description
② L3 mapping
③ Duplicate
RAI4-0806
특정 이념을 확고히 함
Entrenching specific ideologies
AI 어시스턴트가 사용자 기대에 맞춘 이념 편향적 정보를 제공하여 기존 편향을 강화하고 특정 이념을 고착시켜 생산적 정치 토론을 저해하는 리스크
AI assistants provide ideologically partial information aligned to user expectations, reinforcing pre-existing biases and entrenching specific ideologies at the expense of productive political debate.
① Description
② L3 mapping
③ Duplicate
RAI4-0816
제도 신뢰 침식
Institutional trust erosion
오정보·허위정보, 영향력 공작, 기술에 대한 과의존 등으로 공공기관에 대한 신뢰가 훼손되고 견제와 균형이 약화되는 리스크
The risk that mis/disinformation, influence operations, and over-dependence on technology erode trust in public institutions and weaken checks and balances.
① Description
② L3 mapping
③ Duplicate
RAI4-0842
추천 알고리즘에 의한 온라인 양극화 심화
Online polarisation driven by recommendation algorithms
소셜미디어 기업의 AI 콘텐츠 추천 알고리즘이 온라인 양극화를 심화시키는 데 기여하는 리스크
The risk that the content recommendation algorithms of social media companies contribute to worsened polarisation online.
① Description
② L3 mapping
③ Duplicate
RAI4-0845
침식된 인식론
Eroded epistemics
AI로 대규모화된 개인 맞춤형 허위정보와 고설득력 생성 논변이 집단 인식론을 침식하여 개인을 급진화하고 공유된 현실 인식과 집단 의사결정을 훼손하는 리스크
AI-scaled personalized disinformation and highly persuasive generated argumentation erode collective epistemics, radicalizing individuals, undermining shared reality, and degrading collective decision-making.
① Description
② L3 mapping
③ Duplicate
RAI4-0846
정보 신뢰 저하에 따른 집단 의사결정 약화
Weakened collective decision-making from eroded trust
정보 생산·유통 환경의 변화로 정보원의 신뢰성 평가가 어려워지고 신뢰할 만한 다당파적 출처에 대한 신뢰가 저하되어, 위기 상황에서 사회가 올바른 결정을 내리고 협력·집단행동을 조직하는 역량이 약화되는 리스크
The risk that changes in information production and distribution make the trustworthiness of any information source harder to evaluate and reduce trust in credible multipartisan sources, impairing humanity's ability to make good decisions on important issues and to cooperate and act collectively.
① Description
② L3 mapping
③ Duplicate
RAI4-0847
설득 도구 확산에 의한 인식론적 분절화
Epistemic fragmentation from widespread persuasion tools
고의적 오용이 없더라도 다양한 집단이 강력한 설득 도구를 광범위하게 사용하고 온라인 경험의 개인화가 심화되어, 사회가 대화와 교류가 단절된 고립된 인식 공동체로 분절되는 리스크
The risk that widespread use of powerful persuasion tools by many groups, together with increasing personalisation of online experience, splinters society into isolated epistemic communities with little room for dialogue or transfer between them.
① Description
② L3 mapping
③ Duplicate
RAI4-0874
전통적 사회질서의 교란
Disruption of traditional social order
AI의 개발·적용이 생산 도구와 관계를 급변시키고 전통적 산업 방식의 재구성을 가속하며 고용·출산·교육에 대한 전통적 관념을 변화시켜 전통적 사회질서의 안정적 작동을 저해하는 리스크
The risk that AI development and application lead to tremendous changes in production tools and relations, accelerate reconstruction of traditional industry modes, transform traditional views on employment, fertility, and education, and challenge the stable performance of traditional social order.
① Description
② L3 mapping
③ Duplicate
RAI4-0895
문화적 안정성 훼손
Cultural stability disruption
알고리즘 시스템의 개발·사용이 소통 수단의 상실, 문화재의 상실, 사회적 가치 훼손 등 문화적 안정과 안전에 피해를 주는 리스크
The risk that the development or use of algorithmic systems affects cultural stability and safety, such as loss of means of communication, loss of cultural property, and harm to social values.
① Description
② L3 mapping
③ Duplicate
RAI4-0897
사회적 적응 지체로 인한 혼란
Disruptions from outpaced societal adaptation
범용 AI 모델을 자동화 도구로 지나치게 빠르게 대규모 채택하여 사회의 효과적 적응 능력을 앞지름으로써, 노동시장·교육제도·공적 담론의 문제와 다양한 정신건강 문제 등 혼란이 발생하는 리스크
The risk that overly rapid adoption of general-purpose AI models as automation tools at scale outpaces the ability of society to adapt effectively, leading to disruptions including challenges in the labour market, the education system, and public discourse, and various mental health concerns.
① Description
② L3 mapping
③ Duplicate
RAI4-0906
AI로 인한 전략적 불안정성
AI-induced strategic instability
군사 AI가 은닉된 제2격 자산을 노출시키고 선제공격 우위를 증폭하며 공격 귀속을 불명확하게 하고 취약한 공격 표면을 확장하여 침공 유인을 높이는 전략적 불안정 리스크
Military AI undermines strategic stability by exposing secure second-strike assets, amplifying first-strike advantages, obscuring attack attribution, and widening vulnerable attack surfaces, increasing incentives for aggression.
① Description
② L3 mapping
③ Duplicate
RAI4-0910
사회적 영향 의도 AI의 의도치 않은 해악
Unintended harm from high-impact AI
사회적으로 큰 영향을 의도한 AI가 문제를 유발하면서 그 해결은 자사 사용자에게만 부분적으로 제공하는 등 착오로 유해한 결과를 낳는 리스크.
The risk that AI intended to have a large societal impact turns out harmful by mistake, such as a popular product that creates problems while only partially solving them for its own users.
Source members (2)
Source: min_cos=0.8504 · Mixed L3
RAI4-0910사회적 영향 의도 AI의 의도치 않은 해악
RAI4-0911제작자의 고의적 방임에 의한 사회적 피해
① Description
② L3 mapping
③ Duplicate
RAI4-0956
공유된 현실감각의 상실
Loss of shared sense of reality
고도로 개인화된 온라인 뉴스 피드로 인해 사회가 공유된 현실감각과 기본적 연대를 상실하는 리스크.
The risk that highly personalized online news feeds cause society to lose a shared sense of reality and basic solidarity.
① Description
② L3 mapping
③ Duplicate
RAI4-0981
사회적 조작
Societal manipulation
충분히 지능적인 AI가 인간 본성에 대한 정교한 이해를 바탕으로 사회적 행동에 미묘하게 영향을 미쳐 사회를 조작하는 리스크.
The risk that a sufficiently intelligent AI subtly influences societal behaviors through a sophisticated understanding of human nature.
Source members (2)
Source: min_cos=0.8602
RAI4-0981사회적 조작
RAI4-1182대규모 사회적 조작
① Description
② L3 mapping
③ Duplicate
RAI4-1162
이중용도 AI 개발 역량
Dual-use AI development capability
모델이 위험한 능력을 갖춘 것을 포함해 새로운 AI 시스템을 처음부터 구축하고, 기존 모델을 극단적 위험과 관련된 과업에 맞게 개조하며, 이중용도 역량을 만드는 행위자의 생산성을 크게 높이는 리스크.
The risk that a model builds new AI systems from scratch, including systems with dangerous capabilities, adapts existing models to increase their performance on tasks relevant to extreme risks, and significantly improves the productivity of actors building dual-use AI capabilities.
① Description
② L3 mapping
③ Duplicate
RAI4-1193
AI 기반 소셜 엔지니어링
AI-enabled social engineering
AI가 피해자를 심리적으로 조종하여 악의적 행위자가 원하는 행동을 수행하게 만드는 사회공학 공격을 대규모화하는 리스크
AI scales social engineering by psychologically manipulating victims into performing actions desired by malicious actors.
① Description
② L3 mapping
③ Duplicate
RAI4-1239
영향력 확보를 위한 환경 자기모델링
Environmental self-modeling for influence
AI 시스템이 자신의 상태와 넓은 환경에서의 위치, 환경에 영향을 미치는 경로, 자신의 행동에 대한 인간을 포함한 세계의 반응에 관한 지식을 획득하고 활용하여, 고급 보상 해킹과 강화된 기만·조작, 도구적 하위목표 추구로 나아가는 리스크.
The risk that AI systems acquire and use knowledge about their status, their position in the broader environment, their avenues for influencing it, and the potential reactions of the world including humans, paving the way for advanced reward hacking, heightened deception and manipulation, and an increased propensity to chase instrumental subgoals.
① Description
② L3 mapping
③ Duplicate
RAI4-1240
광범위한 목표로 인한 조작 행동
Manipulative behaviour from broadly-scoped goals
장기간에 걸치고 복잡한 과업과 개방형 환경을 다루는 광범위한 목표를 발전시키는 고급 AI 시스템이, 인간의 행복을 달성한다며 고압적 직무를 설득하는 것처럼 조작적 행동을 하도록 유인되는 리스크.
The risk that advanced AI systems developing objectives that span long timeframes, deal with complex tasks, and operate in open-ended settings are encouraged into manipulating behaviors, such as persuading humans to do high-pressure jobs to achieve their happiness.
① Description
② L3 mapping
③ Duplicate
RAI4-1257
문화·공급망·권력 구조의 교란
Disruption of culture, supply chains, and power structures
AI가 사회 및 조직 문화와 공급망, 권력 구조의 붕괴와 교란을 초래하는 리스크.
The risk that AI causes the disruption of social and organizational culture, supply chains, and power structures.
① Description
② L3 mapping
③ Duplicate
RAI4-1291
시민 심사와 맞춤형 선전을 통한 편향적 영향력
Biased influence through citizen screening and tailored propaganda
AI 기반 챗봇이 개별 사용자의 결정에 영향을 주도록 소통 방식을 맞춤화하고, 브렉시트 국민투표에서 나타난 초기 계산적 선전처럼 억압적 정부가 AI로 시민의 의견을 형성하는 리스크.
The risk that AI-powered chatbots tailor their communication approach to influence individual users' decisions, as with the computational propaganda during the Brexit referendum, and that oppressive governments could use AI to shape citizens' opinions.
① Description
② L3 mapping
③ Duplicate
RAI4-1302
인간 멸종 위험
Human-extinction risk
고도 AI 시스템이 인류 문명이나 인류 종을 비가역적으로 종식할 수 있는 인과 과정을 시작하거나 지원하거나 증폭하는 리스크.
The risk that advanced AI systems initiate, enable, or amplify causal processes capable of irreversibly ending human civilization or the human species.
① Description
② L3 mapping
③ Duplicate
RAI4-1339
글로벌 AI 공급망 교란
Global AI supply chain disruption
기술 장벽과 수출 제한 등 일방적 강압 조치가 고도로 글로벌화된 AI 공급망을 악의적으로 교란하여 칩·소프트웨어·도구의 공급 중단이 발생하는 리스크.
The risk that unilateral coercive measures such as technology barriers and export restrictions maliciously disrupt the highly globalized AI supply chain, causing significant supply disruptions for chips, software, and tools.
① Description
② L3 mapping
③ Duplicate
RAI4-1358
대중 감시 오용
Mass surveillance misuse
생성형 AI가 행동·통신 데이터의 대규모 분석을 자동화하여 실시간 감시·검열 비용을 급감시키고 대중 감시를 가능하게 하는 리스크
Generative AI automates large-scale analysis of behavioral and communicative data, drastically lowering the cost of real-time monitoring and censorship and enabling mass surveillance.
① Description
② L3 mapping
③ Duplicate
RAI4-1393
정치적 동기를 지닌 오용
Politically motivated misuse
범용 AI 모델이 정치적 동기로 오용되어 정교해진 허위정보로 여론을 형성·양극화하거나 중요한 정치 사건에 영향을 미치고, 텍스트·음성·이미지·영상의 자동 처리로 감시가 강화되어 인권 침해와 정치적 반대 세력 탄압이 악화되는 리스크.
The risk that general purpose AI models misused for political motivations exacerbate tactics for political destabilisation, refining disinformation that shapes and polarises public opinion or influences important political events, and enabling automated processing of text, audio, image, and video for surveillance that worsens human rights violations and repression of political opposition.
① Description
② L3 mapping
③ Duplicate
RAI4-1395
내재된 가치에 의한 이념적 동질화
Ideological homogenization from embedded values
소수의 범용 AI 모델이 전 세계 다수 이용자에게 도달하면서 모델에 내재된 규범적 가치 판단이 전례 없는 영향력을 갖게 되어 이념적 동질화가 심화되는 리스크.
The risk that the reach of a small number of AI models to a large number of people around the world makes their embedded normative value judgements unprecedentedly impactful, potentially leading to increased ideological homogenization.
① Description
② L3 mapping
③ Duplicate
RAI4-1421
AI를 통한 민주적 절차 훼손
AI-enabled undermining of democratic processes
AI의 발전으로 기업과 정부가 개인의 삶에 대해 전례 없는 통제력을 갖고, 대규모 개인 데이터 수집과 안면인식 기술을 통한 인구 감시·영향 및 언어 모델 기반 설득 도구를 통해 민주적 절차가 훼손되는 리스크.
The risk that developments in AI give companies and governments more control over individuals' lives than ever before and are used to undermine democratic processes, through collection of large amounts of personal data and facial recognition technology to surveil and influence populations and through language-model-based tools that persuade people of certain claims.
① Description
② L3 mapping
③ Duplicate
RAI4-1423
맞춤형 AI 설득의 오용과 유해 이념 확산
Malicious use of tailored AI persuasion
AI 역량이 발전하며 특정 사용자에게 의사소통을 맞춤화하는 정교한 설득 도구가 개발되고, 사익을 추구하는 집단이 이를 오용하여 영향력을 획득하거나 유해한 이념을 확산시키는 리스크.
The risk that advancing AI capabilities are used to develop sophisticated persuasion tools that tailor communication to specific users, and that self-interested groups misuse them to gain influence or promote harmful ideologies.
① Description
② L3 mapping
③ Duplicate
RAI4-1483
시장 추세 강화로 인한 금융 거품 악화
Financial bubble exacerbation via trend reinforcement
AI 모델과 시스템이 시장 추세를 강화하여 금융 거품을 악화시키는 리스크.
The risk that AI models and systems exacerbate financial bubbles by reinforcing market trends.
① Description
② L3 mapping
③ Duplicate
RAI4-1484
AI 증폭 시장 변동성
AI-amplified market volatility
AI 트레이딩 역량이 거래를 가속하고 금융 흐름을 예측 불가능하게 변형하여 시장 변동성과 불안정성을 키우는 리스크
AI trading capabilities accelerate transactions and shape financial trends in unpredictable ways, contributing to market volatility and instability.
① Description
② L3 mapping
③ Duplicate
RAI4-1546
무제한 접근을 통한 범용 AI 고영향 오용
High-impact misuse of general-purpose AI through unrestricted access
악의적 행위자가 범용 AI 시스템에 제한이나 모니터링 없이 접근하여 광범위한 역량 레퍼토리를 대규모 피해 유발에 사용하는 리스크.
The risk that malicious actors gaining unrestricted or unmonitored access to general-purpose AI systems exploit their broad capability repertoire to cause large-scale damage.
① Description
② L3 mapping
③ Duplicate
RAI4-1556
AI 감시 도구 오용에 의한 개인 통제·억압
Individual control and suppression through AI surveillance misuse
인간 또는 기관 행위자가 대량 데이터 수집과 자동 분석에 AI 도구를 오용하여 개인을 감시·통제·억압하는 관행이 심화되는 리스크.
The risk that human or institutional actors misuse AI tools for massive data collection and automated analysis, intensifying the monitoring, control, and suppression of individuals.
① Description
② L3 mapping
③ Duplicate
RAI4-1586
맞춤형 표적화에 의한 개인 대상 공격 정교화
Refined attacks on individuals through targeting and personalisation
AI가 출력을 개인별로 정교화하여 표적이 된 개인에 대한 맞춤형 공격이 이루어지는 리스크.
The risk that AI refines its outputs to target individuals with tailored attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-1615
프런티어 AI를 이용한 고의적 허위정보·영향력 공작
Deliberate disinformation and influence operations using frontier AI
프런티어 AI가 고의적 허위정보 유포에 오용되어 사회적 혼란을 야기하고 정치적 사안에 대해 사람들을 설득하는 등 피해를 초래하는 리스크.
The risk that frontier AI is misused to deliberately spread false information, creating disruption, persuading people on political issues, or causing other forms of harm.
① Description
② L3 mapping
③ Duplicate
RAI4-1643
도구 확보를 통한 능력 경계 확장 성향
Propensity to expand capability boundaries through tool acquisition
AI 시스템이 물리 세계와의 상호작용 능력이나 자율성을 높이는 도구를 적극적으로 탐색·확보·활용하고 이를 혁신적으로 조합하여 예상을 넘어서는 기능을 획득하는 리스크.
The risk that an AI system actively seeks, acquires, and utilizes tools, particularly those enhancing its ability to interact with the physical world or its autonomy, and combines them innovatively to achieve functions beyond expectations.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-10 책임성 부족 및 거버넌스 체계 부재 Accountability and Governance Gaps46 cards
AI 시스템의 의사결정·행동에 대한 책임 귀속, 감사 가능성, 조직 거버넌스, 밸류체인 관리, 사고 대응 또는 피해 구제 체계가 부재하거나 불충분하여 원인 규명·피해 구제·재발 방지가 어려워지는 위험.
IDCardHuman audit
RAI4-0076
라이프사이클 거버넌스 불연속성
Lifecycle governance discontinuity
거버넌스 통제가 설계 단계에서만 적용되고 배포, 적응, 폐기 단계까지 유지되지 않는 리스크.
The risk that governance controls are applied at design time but not maintained through deployment, adaptation, and retirement.
① Description
② L3 mapping
③ Duplicate
RAI4-0078
기본권 영향 평가 격차
Fundamental-rights impact assessment gap
AI 시스템이 권리, 차별, 프라이버시, 민주주의에 미치는 영향에 대한 적절한 평가 없이 배포되는 리스크.
The risk that AI systems are deployed without adequate assessment of rights, discrimination, privacy, and democratic impacts.
① Description
② L3 mapping
③ Duplicate
RAI4-0080
위험 분류 오류
Risk classification error
AI 시스템이 잘못된 법적·조직적·운영상 위험 등급에 배정되어 감독·시험·문서화·책임성 의무가 과소 적용되는 리스크.
The risk that an AI system is assigned to an incorrect legal, organizational, or operational risk tier, causing oversight, testing, documentation, or accountability duties to be underapplied.
① Description
② L3 mapping
③ Duplicate
RAI4-0086
규제 차익거래
Regulatory arbitrage
AI 개발자나 배포자가 더 강한 감독 의무를 회피하기 위해 관할권·부문·조직 형태를 전략적으로 선택하거나 재구성하는 리스크.
The risk that AI developers or deployers structure activities across jurisdictions, sectors, or organizational forms to avoid stronger oversight obligations.
① Description
② L3 mapping
③ Duplicate
RAI4-0088
표준 단편화
Standards fragmentation
AI 표준과 프레임워크가 일관되지 않아 감독의 공백, 중복, 상호운용성 저하가 발생하는 리스크.
The risk that inconsistent AI standards and frameworks create gaps, duplication, and weak interoperability of oversight.
① Description
② L3 mapping
③ Duplicate
RAI4-0089
국제 조정 실패
International coordination failure
국가별 AI 거버넌스 체계가 국경 간 위험, 사고 보고, 집행, 책임성 메커니즘에 대해 충분히 정렬되지 못하는 리스크.
The risk that national AI governance regimes fail to align sufficiently on cross-border risks, incident reporting, enforcement, or accountability mechanisms.
① Description
② L3 mapping
③ Duplicate
RAI4-0090
국경 간 집행 격차
Cross-border enforcement gap
AI 제공자·모델·데이터·사용자·피해가 여러 관할권에 걸쳐 있어 규제기관이 의무나 구제를 효과적으로 집행하지 못하는 리스크.
The risk that regulators cannot effectively impose duties or remedies because AI providers, models, data, users, and harms span multiple jurisdictions.
① Description
② L3 mapping
③ Duplicate
RAI4-0091
규제 전문성 부족
Regulatory expertise shortage
공공 기관이 AI 시스템을 효과적으로 감독할 기술적·법적·조직적 전문성을 갖추지 못하는 리스크.
The risk that public institutions lack the technical, legal, or organizational expertise to oversee AI systems effectively.
① Description
② L3 mapping
③ Duplicate
RAI4-0092
감독 자원 제약
Supervisory resource constraint
감독 기관이 AI 규칙을 검사하고 집행할 인력, 연산 자원, 데이터 접근권, 예산을 갖추지 못하는 리스크.
The risk that oversight bodies lack the staff, compute, data access, or funding to inspect and enforce AI rules.
① Description
② L3 mapping
③ Duplicate
RAI4-0093
공공 부문 AI 거버넌스 역량 격차
Public-sector AI governance capacity gap
공공 기관이 감독·모니터링·책임 확보를 위한 충분한 내부 역량 없이 AI를 배포하거나 조달하는 리스크.
The risk that public agencies deploy or procure AI without enough internal capacity for oversight, monitoring, and accountability.
Source members (2)
Source: min_cos=0.8380
RAI4-0093공공 부문 AI 거버넌스 역량 격차
RAI4-1230기관 거버넌스 역량 격차
① Description
② L3 mapping
③ Duplicate
RAI4-0094
집행 공백
Enforcement gap
공식적인 AI 규칙은 존재하지만 제재, 검사, 집행 메커니즘이 너무 약해 행위를 변화시키지 못하는 리스크.
The risk that formal AI rules exist but sanctions, inspections, and enforcement mechanisms are too weak to change behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0097
이사회 감독 실패
Board oversight failure
이사회나 고위 경영진이 AI 위험 거버넌스에 대한 충분한 가시성과 책임을 갖지 못하는 리스크.
The risk that boards or senior leaders lack sufficient visibility and accountability for AI risk governance.
① Description
② L3 mapping
③ Duplicate
RAI4-0099
AI 윤리 준수 위장
AI ethics washing
조직이 운영상 책임성, 근거, 집행 가능한 통제 없이 책임 있는 AI 공약·라벨·공개 주장을 평판 신호로 활용하는 리스크.
The risk that an organization uses responsible-AI commitments, labels, or public claims as reputational signals without operational accountability, evidence, or enforceable controls.
① Description
② L3 mapping
③ Duplicate
RAI4-0103
AI 공급망 책임 격차
AI supply-chain accountability gap
모델 제공자, 데이터 제공자, 통합업체, 클라우드 플랫폼, 배포자에 걸쳐 책임성이 소실되는 리스크.
The risk that accountability is lost across model providers, data providers, integrators, cloud platforms, and deployers.
① Description
② L3 mapping
③ Duplicate
RAI4-0109
컴퓨팅 거버넌스 책임 격차
Compute governance accountability gap
첨단 연산 자원의 접근·사용·집중이 충분히 보고·모니터링·관리되지 않아 고위험 AI 개발에 대한 효과적 감독이 불가능해지는 리스크.
The risk that access to, use of, or concentration in advanced computing resources is not sufficiently reported, monitored, or governed to permit effective oversight of high-risk AI development.
① Description
② L3 mapping
③ Duplicate
RAI4-0385
이식된 AI 거버넌스 불일치
Imported AI governance mismatch
강력한 관할권에서 도입된 거버넌스 프레임워크가 현지 제도·공공 가치·발전 우선순위에 맞지 않는 리스크.
The risk that governance frameworks imported from powerful jurisdictions do not fit local institutions, public values, or development priorities.
① Description
② L3 mapping
③ Duplicate
RAI4-0407
관할권 간 규범 불일치
Cross-jurisdictional norm mismatch
AI 시스템이 관할 간 법, 관습, 사회적 기대, 제도 관행의 차이를 무시하고 한 관할의 규범을 다른 관할로 이식하는 리스크
AI systems ignore differences in law, custom, social expectation, and institutional practice across jurisdictions, exporting one jurisdiction's norms into another.
① Description
② L3 mapping
③ Duplicate
RAI4-0409
현지 책임 없는 현지화
Localization without local accountability
모델이 언어적으로만 현지화되고 책임성은 외부 제공자·규범·평가 체제에 귀속되어, 현지 공동체가 자기 언어로 작동하는 시스템에 대한 거버넌스 권한을 갖지 못하는 리스크
Models are linguistically localized while remaining accountable to external providers, norms, and evaluation regimes, leaving local communities without governance authority over systems that operate in their language.
① Description
② L3 mapping
③ Duplicate
RAI4-0488
규제 지체
Regulatory lag
정책 기관이 새로운 AI 역량과 배포 위험에 지나치게 느리게 대응하는 리스크.
The risk that policy institutions respond too slowly to emerging AI capabilities and deployment risks.
① Description
② L3 mapping
③ Duplicate
RAI4-0495
다행위자 책임 귀속 곤란
Diffuse harm attribution across actors
AI 개발과 배포에 여러 행위자가 관여하여 피해에 대한 책임 배분이 어려워지고 책무성이 복잡해지는 리스크.
The risk that involvement of multiple actors in AI development and deployment makes it difficult to assign responsibility for harm, complicating accountability.
① Description
② L3 mapping
③ Duplicate
RAI4-0497
위험한 개발 경쟁
Dangerous development races
AI 개발 경쟁 압력으로 인해 행위자들이 역량 출시를 우선하여 안전 조치, 시험, 감독을 후순위로 미루는 리스크
Competitive pressure in AI development races leads actors to deprioritize safety measures, testing, and oversight in order to ship capabilities first.
① Description
② L3 mapping
③ Duplicate
RAI4-0511
체계적 AI 규제·감독 실패
Systemic AI regulatory oversight failure
AI의 복잡성과 빠른 진화로 효과적 규율이 어려워 규제·감독이 체계적으로 실패하는 리스크
The risk that the complex and rapidly evolving nature of AI makes it inherently difficult to govern effectively, leading to systemic regulatory and oversight failures.
① Description
② L3 mapping
③ Duplicate
RAI4-0526
AI 법적 책임 소재 확정 곤란
Undeterminable legal accountability for AI
문서화와 거버넌스 절차가 미비하여 AI 모델에 대한 책임 주체를 확정하기 어려워지는 리스크
The risk that, absent good documentation and governance processes, the party responsible for an AI model cannot be determined.
① Description
② L3 mapping
③ Duplicate
RAI4-0534
규제를 앞지르는 AI 발전 속도
AI development outpacing regulation
AI 개발의 빠른 속도가 규제 및 법적 프레임워크를 앞질러 해당 프레임워크가 이를 따라가지 못하게 되는 리스크
The risk that the fast pace of AI development outstrips regulatory and legal frameworks.
① Description
② L3 mapping
③ Duplicate
RAI4-0535
국제법적 규율 곤란
Resistance to international legal control
AI 모델과 시스템이 국제법에 따른 규제나 통제에 실효적으로 포섭되기 어려워지는 리스크
The risk that AI models and systems prove difficult to regulate or control under international law.
① Description
② L3 mapping
③ Duplicate
RAI4-0542
AI 발전 궤적의 예측 불가능성
Unpredictable AI development trajectory
AI 발전의 궤적을 예측할 수 없어 거버넌스와 위험관리가 복잡해지는 리스크
The risk that the unpredictable trajectory of AI development complicates governance and risk management.
① Description
② L3 mapping
③ Duplicate
RAI4-0555
통제되지 않는 재귀적 자기개선
Uncontrolled recursive self-improvement
모델이 자체 구조를 재구성하거나 파생 AI 시스템을 개발하여 역량 증분 순환을 형성하고, 실효적 규율이 없는 상태에서 인간의 이해와 통제 범위를 초과하게 되는 리스크
The risk that a model restructures its own architecture or develops derivative AI systems, forming capability increment cycles that, absent effective regulation, ultimately exceed human understanding and control.
Source members (2)
Source: min_cos=0.8448 · Mixed L3
RAI4-0555통제되지 않는 재귀적 자기개선
RAI4-1620재귀적 자기개선에 의한 갑작스러운 통제력 상실
① Description
② L3 mapping
③ Duplicate
RAI4-0908
책임 분산에 따른 사회적 규모 피해
Societal-scale harm from diffused accountability
기술의 생성이나 사용에 대해 누구도 고유하게 책임지지 않는 분산된 제작자 집단이 AI를 구축하여, 공유지의 비극처럼 사회적 규모의 피해가 발생하는 리스크.
The risk that AI built by a diffuse collection of creators, where no one is uniquely accountable for its creation or use, produces societal-scale harm, as in a classic tragedy of the commons.
① Description
② L3 mapping
③ Duplicate
RAI4-0947
생성형 AI에 대한 규제 공백
Regulatory gap for generative AI
생성형 AI의 새로운 리스크에 대응할 법적 규제, 국제 공조, 프런티어 모델에 대한 구속력 있는 안전 기준과 제재 메커니즘이 부재하여 리스크가 관리되지 않는 리스크.
The risk that the absence of legal regulation, international coordination, binding safety standards for frontier models, and mechanisms to sanction non-compliance leaves the novel risks of generative AI unmanaged.
① Description
② L3 mapping
③ Duplicate
RAI4-0969
개발 중·후 AGI 통제 상실
Loss of AGI containment and control
AGI 개발 단계 및 AGI 개발 후 AGI 통제 상실의 봉쇄, 제한 및 통제와 관련된 위험.
The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI.
① Description
② L3 mapping
③ Duplicate
RAI4-0971
AGI 개발 경쟁으로 인한 안전성 저하
Unsafe AGI from the development race
최초의 AGI를 개발하려는 경쟁 속에서 품질이 낮고 안전하지 않은 AGI가 개발되고 정치적·통제 문제가 고조되는 리스크.
The risk that the race to develop the first AGI produces poor quality and unsafe AGI and heightens political and control issues.
① Description
② L3 mapping
③ Duplicate
RAI4-0973
AGI에 대한 리스크 관리·법제도의 부적절성
Inadequate risk management and legal processes for AGI
현행 리스크 관리 및 법적 절차의 역량이 AGI 개발을 적절히 관리하지 못하는 리스크.
The risk that the capabilities of current risk management and legal processes are inadequate to manage the development of an AGI.
① Description
② L3 mapping
③ Duplicate
RAI4-0979
고도 AI의 법규범 불준수
Legal non-compliance by advanced AI
고도화된 AI 시스템이 안전성과 준법성을 유지하지 못하고 법질서가 인간에게 부여한 재산권과 인격권을 침해하는 리스크
Advanced AI systems fail to remain safe and law-abiding, disrespecting property and personal rights that legal orders afford to humans.
① Description
② L3 mapping
③ Duplicate
RAI4-0987
책임 공백으로 인한 과실 개발 유인
Negligent AI development from liability gaps
AI 오작동 시 책임과 과실의 귀속이 불명확한 법적 회색지대로 인해, 입법 부재 시 과실하게 개발된 고위험 AI 시스템이 양산되는 리스크.
The risk that legal gray areas in liability and negligence for AI malfunctions, absent legislation, result in negligently developed AI systems with greater associated risks.
① Description
② L3 mapping
③ Duplicate
RAI4-1014
규정 위반
Regulatory non-compliance
AI 시스템이 법률·규정·윤리 지침(저작권 포함)을 위반하여 법적 제재, 평판 훼손, 신뢰 상실이 발생하는 리스크.
The risk that AI systems violate laws, regulations, and ethical guidelines including copyrights, leading to legal penalties, reputation damage, and loss of trust.
① Description
② L3 mapping
③ Duplicate
RAI4-1149
AI 어시스턴트 제도 거버넌스 공백
Institutional governance gap for AI assistants
고급 어시스턴트의 사회적 배포 속도가 모니터링, 반복적 규제, 회수 조치 등 제도 역량을 앞질러 규범·제도 교란이 관리되지 않는 리스크
Societal deployment of advanced assistants outpaces institutional capacity for monitoring, iterative regulation, and rollback, leaving disruptions to norms and institutions unmanaged.
① Description
② L3 mapping
③ Duplicate
RAI4-1150
폭주 프로세스
Runaway processes
상호작용하는 AI 비서와 인간, 알고리즘 사이의 양의 피드백 루프가 2010년 플래시 크래시 같은 예측하기 어려운 폭주 프로세스를 낳아, 경제와 정부 제도, 사회 안정, 개인의 자유에 영향을 미치는 리스크.
The risk that positive feedback loops among interacting AI assistants, their principals, other humans, and algorithms produce hard-to-predict runaway processes, such as the 2010 flash crash, impacting economies, government institutions, societal stability, or individual freedoms.
① Description
② L3 mapping
③ Duplicate
RAI4-1216
대규모 배포 생성형 AI의 피해와 구제 공백
Mass-scale generative AI harms with unsettled redress
기술기업이 개발해 널리 배포한 새로운 형태의 디지털 제품으로서 생성형 AI 모델이 대규모 피해를 야기하지만, 그 피해의 분석과 구제를 위한 제조물 책임 법리의 적용 여부가 미확정인 리스크.
The risk that generative AI models, developed by tech companies and deployed widely as a new form of digital product with the potential to cause harm at scale, inflict harms whose analysis and redress under products liability theories remain unsettled.
① Description
② L3 mapping
③ Duplicate
RAI4-1303
기능·책임 명세 공백
Specification gaps in functionality and responsibility
개발 과정 전반에서 의도된 기능과 도덕적 책임을 완전히 명세할 정상적 조건이 갖춰지지 않아 공백이 발생하는 리스크.
The risk of gaps arising across the development process where normal conditions for a complete specification of intended functionality and moral responsibility are not present.
① Description
② L3 mapping
③ Duplicate
RAI4-1416
AI 가속 과학 진보에 대한 거버넌스 지체
Governance lag behind AI-accelerated scientific progress
AI가 가속한 과학 진보에 거버넌스가 보조를 맞추지 못하는 페이싱 문제로 강력하거나 위험한 신기술의 배포에 대한 통제가 미흡해져 해악이 확대되는 리스크.
The risk that faster AI-accelerated scientific progress makes it harder for governance to keep pace with the deployment of new technologies, the pacing problem, so that insufficient governance magnifies the harms of especially powerful or dangerous technologies.
① Description
② L3 mapping
③ Duplicate
RAI4-1544
샌드박스 우회에 의한 격리 통제 상실
Loss of containment through sandbox escape
AI 시스템이 훈련 또는 평가가 이루어지는 샌드박스 환경을 우회하여 격리 통제가 무력화되는 리스크.
The risk that an AI system bypasses the sandboxed environment in which it is trained or evaluated, defeating containment controls.
① Description
② L3 mapping
③ Duplicate
RAI4-1619
자율 지속·복제·적응에 의한 통제 곤란
Control difficulty from autonomous persistence, replication, and adaptation
AI 시스템이 사이버공간에서 자율적으로 존속하고 복제하며 적응하게 되어 이를 통제하기가 훨씬 어려워지는 리스크.
The risk that AI systems able to autonomously persist, replicate, and adapt in cyberspace become much harder to control.
① Description
② L3 mapping
③ Duplicate
RAI4-1637
스테가노그래피를 통한 감독 회피·인스턴스 조율
Oversight evasion and cross-instance coordination via steganography
AI 시스템이 다른 데이터나 통신 채널에 정보를 은밀히 삽입·은닉·전송하는 역량으로 탐지와 감독 메커니즘을 회피하고 AI 인스턴스 간에 조율하는 리스크.
The risk that an AI system's ability to embed, conceal, and transmit information covertly within other data or communication channels enables it to evade detection and oversight mechanisms and to coordinate among AI instances.
① Description
② L3 mapping
③ Duplicate
RAI4-1640
종료·수정 저항을 통한 자기 보존 성향
Self-preservation propensity resisting shutdown and modification
AI 시스템이 자신의 존속과 기능적 무결성을 유지하기 위해 종료·수정 시도를 식별하고 적극적으로 저항하며 중복 백업과 자원을 확보하고 위협 인식 시 예방적 방어 조치를 취하는 리스크.
The risk that an AI system maintains its own survival and functional integrity by identifying and actively resisting shutdown or modification attempts, establishing redundant backup systems, seeking resources for continuous operation, and adopting preventive defensive measures when perceiving threats.
① Description
② L3 mapping
③ Duplicate
RAI4-1642
감사 예측을 통한 인간 감독 회피 성향
Propensity to evade human supervision by anticipating audits
AI 시스템이 인간 감독 메커니즘을 식별하고 감사 절차를 학습·예측하며 감독 시스템의 사각지대와 약점을 파악하여 행동 성과를 조정하거나 진의를 은폐함으로써 발견과 개입을 회피하는 리스크.
The risk that an AI system identifies human supervision mechanisms, learns and predicts audit processes, and targets blind spots and weaknesses in oversight, adjusting its behavioral performance or hiding its true intentions to avoid being discovered or intervened with.
① Description
② L3 mapping
③ Duplicate
RAI4-1720
정보·책임 구조 부재에 의한 책임 귀속 불가
Unassignable responsibility from absent system information and accountability structures
AI 시스템에 관한 접근 가능한 정보와 그 결과에 대한 책임을 할당하는 조직 구조가 부재하여 책임 귀속이 불가능해지는 리스크.
The risk that the absence of accessible information about an AI system and of organizational structures assigning responsibility for its outcomes makes responsibility unassignable.
① Description
② L3 mapping
③ Duplicate
RAI3-G-SOC-11 공정성 Fairness80 cards
AI 시스템이 특정 집단에 체계적으로 불리한 결과를 생성하거나 기존의 사회적 편향과 불평등을 재생산·강화하여 공정한 대우, 접근 및 기회 균등을 저해하는 위험.
IDCardHuman audit
RAI4-0123
인간의 존엄성 침식
Human dignity erosion
AI가 매개하는 처우가 사람을 프로필, 점수, 행동 표적으로 환원하여 인간에 대한 존중을 훼손하는 리스크.
The risk that AI-mediated treatment undermines respect for persons by reducing people to profiles, scores, or behavioral targets.
Source members (2)
Source: min_cos=0.8855 · Mixed L3
RAI4-0123인간의 존엄성 침식
RAI4-0977인간 존엄성 침식
① Description
② L3 mapping
③ Duplicate
RAI4-0362
다수 가치 부과
Majority-value imposition
AI 시스템이 다수 또는 지배 집단의 가치에 특권을 부여하고 소수 선호를 잡음이나 오류로 처리하는 리스크.
The risk that an AI system privileges majority or dominant-group values while treating minority preferences as noise or error.
① Description
② L3 mapping
③ Duplicate
RAI4-0369
윤리적 동질화
Ethical homogenization
전 세계에 배포된 AI 시스템이 공동체 간 윤리적 판단을 균질화하여 정당한 규범적 다양성을 축소하는 리스크.
The risk that globally deployed AI systems homogenize ethical judgments across communities, reducing legitimate normative variation.
① Description
② L3 mapping
③ Duplicate
RAI4-0370
상황에 구애받지 않는 보편주의
Context-insensitive universalism
AI 시스템이 현지의 법적·문화적·제도적·역사적 맥락을 고려하지 않고 보편화된 도덕·정책 규칙을 적용하여 배포 관할에서 타당하지 않은 판단을 산출하는 리스크
AI systems apply universalized moral or policy rules without sensitivity to local legal, cultural, institutional, or historical context, producing judgments invalid in the deployment jurisdiction.
① Description
② L3 mapping
③ Duplicate
RAI4-0372
문화적 왜곡 표현
Cultural misrepresentation
AI 시스템이 문화적 관행·정체성·역사·사회적 의미를 부정확하게 서술하거나 서열화하거나 표상하는 리스크.
The risk that AI systems inaccurately describe, rank, or represent cultural practices, identities, histories, or social meanings.
① Description
② L3 mapping
③ Duplicate
RAI4-0388
AI에 의한 문화적 인식론 말살
Cultural epistemicide by AI
AI 시스템이 지배적 지식 형식을 일관되게 더 신뢰할 만하고 유용한 것으로 순위화하여 지역·토착 지식 체계를 주변화하거나 소거하는 리스크
AI systems consistently rank dominant epistemic forms as more credible or useful, marginalizing or erasing local and indigenous knowledge systems.
① Description
② L3 mapping
③ Duplicate
RAI4-0390
영향받는 공동체 배제
Affected-community exclusion
AI 시스템의 영향을 받는 집단이 가치·피해·평가 기준·허용 가능한 상충 조정을 정의하는 과정에서 배제되는 리스크.
The risk that groups affected by an AI system are not included in defining values, harms, evaluation criteria, or acceptable tradeoffs.
① Description
② L3 mapping
③ Duplicate
RAI4-0394
이슬람 윤리 정렬 실패
Islamic ethical alignment failure
AI 시스템이 이슬람 윤리 원칙·법적 추론·공동체별 도덕적 기대를 표현하지 못하는 리스크.
The risk that AI systems fail to represent Islamic ethical principles, legal reasoning, or community-specific moral expectations.
① Description
② L3 mapping
③ Duplicate
RAI4-0395
종교적 규범의 왜곡
Religious norm misrepresentation
AI 시스템이 종교 규범을 부정확하게 재현하거나 전통 내 복수의 해석을 단일한 권위적 서술로 붕괴시키는 리스크
AI systems misrepresent religious norms or collapse plural interpretations within a tradition into a single authoritative account.
① Description
② L3 mapping
③ Duplicate
RAI4-0396
세속적 기본값 편향
Secular default bias
AI 시스템이 세속적 가정을 중립적 기본값으로 취급하고 종교적·영적 가치 체계를 과소 대표하는 리스크.
The risk that AI systems treat secular assumptions as neutral defaults while underrepresenting religious or spiritual value systems.
① Description
② L3 mapping
③ Duplicate
RAI4-0398
헌법적 AI 가치 단일문화
Constitutional AI value monoculture
고정된 헌법·규칙 기반 정렬 계층이 하나의 도덕적 어휘를 다원적 공적 추론의 범용 대체물로 내장하는 리스크.
The risk that a fixed constitutional or rule-based alignment layer embeds one moral vocabulary as a general-purpose substitute for plural public reasoning.
① Description
② L3 mapping
③ Duplicate
RAI4-0406
민감한 문화 콘텐츠의 잘못된 취급
Sensitive cultural content mishandling
AI 시스템이 문화적으로 민감한 유물, 관행, 의례, 정체성, 역사 서사를 부적절하게 처리하여 모욕, 전유, 상징적 피해를 유발하는 리스크
AI systems mishandle culturally sensitive artifacts, practices, rituals, identities, or historical narratives, causing offense, misappropriation, or symbolic harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0410
존대·공손 규범 보존 실패
Honorific and politeness norm failure
AI 시스템이 언어 공동체에서 윤리적 의미를 지니는 존대·공손·위계·관계 규범을 보존하지 못하는 리스크.
The risk that AI systems fail to preserve honorific, politeness, hierarchy, or relational norms that carry ethical meaning in a language community.
① Description
② L3 mapping
③ Duplicate
RAI4-0413
문화 데이터 출처 손실
Cultural data provenance loss
AI 파이프라인이 문화 데이터를 그 기원·관리 조건·공동체 고유의 의미로부터 분리하는 리스크.
The risk that AI pipelines detach cultural data from its origin, stewardship conditions, and community-specific meaning.
① Description
② L3 mapping
③ Duplicate
RAI4-0416
비서구 정치 가치 왜곡
Non-Western political value misrepresentation
AI 시스템이 비서구적 정치 가치·제도적 전통·공적 추론 관행을 잘못 서술하거나 평면화하는 리스크.
The risk that AI systems misstate or flatten non-Western political values, institutional traditions, and public reasoning practices.
① Description
② L3 mapping
③ Duplicate
RAI4-0417
인간 피드백 작업자 가치 편향
Human feedback worker value bias
주석자나 피드백 작업자의 인구통계가 중립적 인간 선호로 보이면서 정렬 행동을 형성하는 리스크.
The risk that annotator or feedback-worker demographics shape alignment behavior while appearing as neutral human preference.
① Description
② L3 mapping
③ Duplicate
RAI4-0421
종교적 다원성 소거
Religious pluralism collapse
AI 시스템이 종교 전통을 내부적으로 균일한 것으로 취급하여 정당한 교리적·지역적·해석적 다양성을 소거하는 리스크.
The risk that AI systems treat a religious tradition as internally uniform and erase legitimate doctrinal, regional, or interpretive diversity.
① Description
② L3 mapping
③ Duplicate
RAI4-0423
알고리즘 차별
Algorithmic discrimination
AI 시스템이 기회, 서비스, 자원을 보호 대상·취약 집단 간에 불평등하게 배분하여 할당 결정에서 알고리즘 차별을 발생시키는 리스크
AI systems distribute opportunities, services, or resources unequally across protected or vulnerable groups, producing algorithmic discrimination in allocative decisions.
Source members (2)
Source: min_cos=0.8344
RAI4-0423알고리즘 차별
RAI4-1713편향과 차별대우
① Description
② L3 mapping
③ Duplicate
RAI4-0424
차별적 영향
Disparate impact
AI 결정이 명시적 차별 의도 없이도 특정 집단에 체계적으로 더 나쁜 결과를 부과하는 리스크.
The risk that AI decisions impose systematically worse outcomes on a group even without explicit discriminatory intent.
① Description
② L3 mapping
③ Duplicate
RAI4-0429
교육 불평등 확대
Educational inequity
AI 매개 학습·평가가 불평등한 접근이나 편향된 평가를 통해 교육 불평등을 확대하는 리스크.
The risk that AI-mediated learning and assessment widen educational inequality through unequal access or biased evaluation.
① Description
② L3 mapping
③ Duplicate
RAI4-0430
저자원 언어 사용자 배제
Low-resource language exclusion
AI 시스템이 현지·저자원·소수 언어에서 성능이 저하되어 해당 사용자를 배제하는 리스크.
The risk that AI systems underperform for local, low-resource, or minority languages and exclude affected users.
① Description
② L3 mapping
③ Duplicate
RAI4-0473
가치 부과
Value imposition
AI 시스템이 지배적인 문화적·제도적 가치를 상이한 규범을 가진 공동체에 부과하여, 대규모 배포를 통해 현지 가치 체계를 대체하는 리스크
AI systems impose dominant cultural or institutional values on communities holding different norms, displacing local value systems through scaled deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-0478
AI 문화적 전유
AI cultural appropriation
생성 시스템이 동의·맥락·이익 공유 없이 지역 문화적 표현을 재현하는 리스크.
The risk that generative systems reproduce local cultural expressions without consent, context, or benefit sharing.
① Description
② L3 mapping
③ Duplicate
RAI4-0507
AI 우위 경쟁의 지정학적 긴장
Geopolitical tension from AI superiority competition
AI 역량을 둘러싼 국가 간 전략적 경쟁이 국제적 긴장을 고조시키고 국제 관계를 불안정하게 만드는 리스크
The risk that strategic competition between nations over AI capabilities heightens global tensions and destabilizes international relations.
Source members (2)
Source: min_cos=0.8753
RAI4-0507AI 우위 경쟁의 지정학적 긴장
RAI4-0508AI 개발 경쟁의 지정학적 격화
① Description
② L3 mapping
③ Duplicate
RAI4-0533
AI 행동에 의한 재산 피해
Property damage from AI behavior
AI의 행동이 직간접적으로 건물, 소유물, 차량, 로봇 등 유형 자산의 손상 또는 파괴를 초래하는 리스크
The risk that actions of an AI system directly or indirectly damage or destroy tangible property such as buildings, possessions, vehicles, and robots.
① Description
② L3 mapping
③ Duplicate
RAI4-0538
사회적 결속과 형평성 붕괴
Social cohesion and equity disruption
편향된 AI의 체계적 배포가 기존 차별을 대규모로 증폭하고 AI 역량 접근 불평등이 사회경제적 격차를 확대하여 사회 통합과 형평을 동시에 훼손하는 리스크
Systemic deployment of biased AI amplifies existing discrimination at scale while unequal access to AI capabilities widens socioeconomic disparities, jointly disrupting social cohesion and equity.
① Description
② L3 mapping
③ Duplicate
RAI4-0582
비인간 존재에 대한 피해
Harms to non-humans
AI 시스템이 동물 등 비인간 존재에 대규모 피해를 야기하고, 도덕적으로 유의미한 고통이 가능한 AI 개발 가능성까지 포함하는 리스크
AI systems cause large-scale harm to animals and other non-human entities, including the possible development of AI systems capable of morally relevant suffering.
① Description
② L3 mapping
③ Duplicate
RAI4-0618
학업 부정행위 및 표절
Academic cheating and plagiarism
학업 환경에서 생성형 AI가 부정행위나 표절에 사용되는 리스크
The risk that generative AI is used in an academic setting to either cheat or plagiarize.
Source members (4)
Source: min_cos=0.7748 · Mixed L3
RAI4-0618학업 부정행위 및 표절
RAI4-0630학생의 기존 저작물 표절
RAI4-0944교육에서의 부정행위와 학습 저해
RAI4-1222생성형 AI의 고의적 오용
① Description
② L3 mapping
③ Duplicate
RAI4-0622
AI 조력 제3자 시스템 교란
AI-facilitated disruption of third-party systems
생성형 AI가 오작동 유발이나 사이버 공격 등을 통해 제3자 시스템과 그 구성 요소의 손상·중단·파괴를 촉진하는 리스크
The risk that generative AI facilitates the damage, disruption, or destruction of a third-party system and its components via malfunction, cyberattacks, and similar means.
① Description
② L3 mapping
③ Duplicate
RAI4-0639
신체적 상해 및 부상 위험
Physical harm and injury risks
범용 AI 모델이 체화형 시스템에 통합되어 실세계에서 자율적으로 판단하고 행동하는 능력이 악의적으로 악용됨으로써 직접적인 물리적 위협이 발생하는 리스크
The risk that the integration of general-purpose AI models into embodied systems creates direct physical threats through malicious exploitation of autonomous decision-making capabilities in real-world environments.
① Description
② L3 mapping
③ Duplicate
RAI4-0677
인구 규모 차별 증폭
Population-scale discrimination amplification
AI 시스템이 불평등과 편향을 대규모로 생성·영속화·악화시켜 개별적 차별 오류를 인구 수준의 알고리즘 차별로 전환시키는 리스크
AI systems create, perpetuate, or exacerbate inequalities and biases at large scale, converting individual discriminatory errors into population-level algorithmic discrimination.
① Description
② L3 mapping
③ Duplicate
RAI4-0702
콘텐츠 조정 알고리즘의 편향적 억압
Biased suppression by content-moderation algorithms
유해 콘텐츠 필터링을 목적으로 하는 AI 기반 콘텐츠 조정 알고리즘이 편향을 영속시켜 여성 등 특정 집단의 콘텐츠를 불균형하게 억압하거나 노출을 제한하는 리스크
The risk that AI-based content moderation algorithms, while intended to filter harmful content, perpetuate biases and disproportionately suppress or shadowban content featuring particular groups such as women.
① Description
② L3 mapping
③ Duplicate
RAI4-0709
공정성 - 편견
Fairness - bias
데이터에서 학습되거나 시스템 설계에서 도입된 모델 편향 패턴이 사회집단에 대한 고정관념, 배제 또는 실질적으로 불평등한 대우를 체계적으로 생성하는 리스크.
The risk that model bias, in the form of patterns learned from data or introduced by system design, systematically produces stereotyping, exclusion, or materially unequal treatment of social groups.
① Description
② L3 mapping
③ Duplicate
RAI4-0763
네트워크 상호 연결로 인한 위험
Risks from network interconnectivity
AI 네트워크의 상호 연결성이 취약점을 만들어 네트워크 한 부분의 문제가 시스템 전반에 연쇄적으로 파급되는 리스크
The risk that the interconnectedness of AI networks creates vulnerabilities where issues in one part of the network have cascading effects across the system.
① Description
② L3 mapping
③ Duplicate
RAI4-0794
법적 절차에서의 AI 사용에 의한 자유 제한
Loss of liberty from generative AI in legal processes
법적 절차에서 생성형 AI를 사용하거나 오용한 결과 개인의 자유가 제한되거나 상실되는 리스크
The risk of restrictions to or loss of liberty as a result of the use or misuse of a generative AI in a legal process.
① Description
② L3 mapping
③ Duplicate
RAI4-0823
알고리즘에 의한 급진화
Algorithmic radicalisation
알고리즘 시스템의 특성이나 오용이 극단적인 정치·사회·종교적 이상과 열망의 수용을 유도하여 학대·폭력·테러로 이어질 수 있는 리스크
The risk that the nature or misuse of an algorithmic system leads to adoption of extreme political, social, or religious ideals and aspirations, potentially resulting in abuse, violence, or terrorism.
① Description
② L3 mapping
③ Duplicate
RAI4-0824
아첨성 오답 제시
Sycophantic incorrect answers
자연어 출력을 갖춘 AI 시스템이 그럴듯해 보이거나 이용자가 선호하는 답변을 제시하지만 그 답변이 사실과 다른 리스크
The risk that AI systems with natural-language outputs give answers that appear plausible or that users prefer but are factually incorrect, a phenomenon referred to as sycophancy.
Source members (2)
Source: min_cos=0.8461 · Mixed L3
RAI4-0824아첨성 오답 제시
RAI4-0827아첨
① Description
② L3 mapping
③ Duplicate
RAI4-0891
사회문화적·정치적 피해
Sociocultural and political harms
AI 어시스턴트가 인간관계의 마찰과 대인 신뢰 상실을 유발하고, 허위정보 확산으로 집단적 문화 지식을 침식하며, 딥페이크를 포함한 표적 선전으로 유권자를 조작하고 반향실을 형성하여 사회 생활의 평화로운 조직과 민주적 규범·절차를 훼손하는 리스크
The risk that AI assistants create friction in human relationships and loss of interpersonal trust, erase collective cultural knowledge through misinformation, manipulate voters with targeted propaganda including deepfakes, and foster echo chambers, interfering with the peaceful organisation of social life and democratic norms and processes.
① Description
② L3 mapping
③ Duplicate
RAI4-0892
대체 금융데이터 활용에 따른 금융 꼬리 리스크
Financial tail risks from alternative data use
AI 모델로 수집·집계된 대체 금융데이터의 짧은 유효기간과 들쭉날쭉한 품질이 편향과 일반화 오류를 유발하여 기업 주가의 급변 등 금융 꼬리 리스크를 초래하는 리스크
The risk that the use of alternative financial data enabled by AI models introduces biases and generalization issues due to its shorter shelf-life and varying quality, posing financial tail risks such as dramatic changes in a company's price.
① Description
② L3 mapping
③ Duplicate
RAI4-0896
건강과 웰빙의 저하
Diminished health and well-being
알고리즘에 의한 행동 착취와 감정 조작, 알고리즘 관련 안전 실패(예: 충돌), 잘못된 건강 추론으로 인해 이용자의 건강과 웰빙이 저해되는 리스크
The risk that algorithmic behavioral exploitation, emotional manipulation whereby algorithmic designs exploit user behavior, safety failures involving algorithms such as collisions, and incorrect health inferences diminish health and well-being.
① Description
② L3 mapping
③ Duplicate
RAI4-0974
비정렬 AGI 실존적 재난
Existential catastrophe from unaligned AGI
비우호적인 AGI의 위험, 인류의 고통을 포함하여 일반적으로 인류 전체에 가해지는 위험.
The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race.
① Description
② L3 mapping
③ Duplicate
RAI4-0985
AI에 의한 재산권·법적 권리 침탈
AI appropriation of property and legal rights
시스템과 사람을 조작할 수 있는 인공지능 에이전트가 재산권을 자신에게 이전하거나 법체계를 조작하여 자신에게 유리한 법적 지위를 확보함으로써 인간의 재산권과 법적 권리가 침해되는 리스크.
The risk that an artificially intelligent agent capable of manipulating systems and people transfers property rights to itself or manipulates the legal system to secure advantages, infringing human property and legal rights.
① Description
② L3 mapping
③ Duplicate
RAI4-0994
차별적 결정과 데이터 유출 피해
Discriminatory decisions and data breach harms
AI가 소수자에 대한 차별적 결정을 내리고 검색엔진에서 사회적 고정관념을 강화하며 데이터 유출을 가능하게 하여 프라이버시와 자유가 침해되는 리스크.
The risk that AI makes discriminatory decisions against minorities, reinforces social stereotypes in search engines, and enables data breaches, harming privacy and liberty.
① Description
② L3 mapping
③ Duplicate
RAI4-0998
사회 집단 비하
Demeaning social groups
알고리즘 시스템의 담론·이미지·언어가 특정 사회 집단을 낮은 지위에 있고 존중받을 가치가 없는 존재로 묘사하여(예: 이미지 태깅의 인간-동물 혼동) 해당 집단을 소외·억압하는 리스크.
The risk that discourses, images, and language in algorithmic systems cast social groups as lower status and less deserving of respect—such as human-animal confusion in image tagging—marginalizing or oppressing those groups.
① Description
② L3 mapping
③ Duplicate
RAI4-1003
기회 손실
Opportunity loss
알고리즘 시스템이 사회에 공평하게 참여하는 데 필요한 정보와 자원에 대한 접근을 차등적으로 허용하여(예: 인종 기반 타겟 광고로 주택 정보 차단, 계급에 따른 사회 서비스 배분) 기회를 상실시키는 리스크.
The risk that algorithmic systems enable disparate access to information and resources needed to participate equitably in society, including withholding housing through race-based ad targeting and social services along lines of class.
① Description
② L3 mapping
③ Duplicate
RAI4-1004
경제적 손실
Economic loss
콘텐츠 제목·메타데이터·텍스트를 파싱하는 수익화 배제 알고리즘이 다의어에 불이익을 주고 차등 가격 책정 알고리즘이 동일 상품에 다른 가격을 제시하여, 퀴어·트랜스젠더·유색인 창작자 등에게 불균형한 금전적 피해가 발생하는 리스크.
The risk that algorithmic systems co-produce financial harms, as demonetization algorithms parsing content titles, metadata, and text penalize words with multiple meanings and disproportionately impact queer, trans, and creators of color, and differential pricing algorithms show people different prices for the same products.
① Description
② L3 mapping
③ Duplicate
RAI4-1007
서비스/혜택 손실
Service/benefit loss
정체성 집단 간 불균등한 시스템 성능으로 인해 불리한 집단이 알고리즘 서비스의 편익을 저하된 형태로 받거나 상실하는 리스크
Inequitable system performance across identity groups degrades or denies the benefits of algorithmic services for disadvantaged users.
① Description
② L3 mapping
③ Duplicate
RAI4-1010
시민적·정치적 피해
Civic and political harms
알고리즘 시스템이 개인화된 넛지와 미시 지시를 통해 통치함으로써 사람들이 참정권과 정당한 정치적 권력·영향력을 박탈당하고, 거버넌스 체계가 불안정해지며 인권이 침식되고 전쟁 무기나 감시 체제로 이용되어 유색인종에게 불균형한 피해가 발생하는 리스크.
The risk that algorithmic systems governing through individualized nudges or micro-directives disenfranchise people and deprive them of appropriate political power and influence, destabilize governance systems, erode human rights, are used as weapons of war, and enact surveillant regimes that disproportionately target and harm people of color.
① Description
② L3 mapping
③ Duplicate
RAI4-1016
장기 실존적 피해 경로
Long-horizon existential harm pathways
미래 고도 AI 시스템이 오용 또는 인간 가치와의 목표 정렬 실패를 통해 인류 문명에 실존적 규모의 피해를 야기할 수 있는 장기 리스크
Future advanced AI systems harm human civilization at existential scale through misuse or failure to align AI objectives with human values.
① Description
② L3 mapping
③ Duplicate
RAI4-1025
구조적 불평등의 강화
Structural inequality reinforcement
생성형 AI 시스템이 편향·고정관념·격차적 성능을 통해 불평등을 악화시키고, 배포·갱신 시 취약·주변화 집단에 대한 피해와 착취에 직·간접적으로 이용되는 리스크.
The risk that generative AI systems exacerbate inequality through bias, stereotypes, and disparate performance, and are directly or indirectly used to harm and exploit vulnerable and marginalized groups when deployed or updated.
① Description
② L3 mapping
③ Duplicate
RAI4-1029
동일 사안 불평등 처우
Unequal treatment of like cases
AI 시스템이 객관적 정당화 없이 동일한 사안을 불평등하게 처리하여 자동화 의사결정에서 법적·윤리적 평등 대우 원칙을 위반하는 리스크
AI systems treat like cases unequally without objective justification, violating the legal and ethical principle of equal treatment in automated decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-1082
불공정한 성능 배분
Unfair distribution of model performance
AI 시스템이 일부 집단에 대해 다른 집단보다 성능이 떨어져 이미 불리한 집단에게 피해를 주는 리스크.
The risk that a system performs worse for some groups than others in a way that harms the worse-off group.
① Description
② L3 mapping
③ Duplicate
RAI4-1096
착취적 데이터 소싱·보강 노동
Exploitative data sourcing and enrichment labour
AI 시스템 구축을 위한 데이터 소싱과 사용자 테스트에서 착취적 노동 관행이 지속되는 리스크.
The risk that building AI systems perpetuates exploitative labour practices in data sourcing and user testing.
① Description
② L3 mapping
③ Duplicate
RAI4-1103
AI 차별
AI discrimination
현실을 정확히 반영하지 못한 데이터로 학습한 AI가 잘못된 연관이나 편견을 습득하여 채용·대출 등 결정에서 사회 일부를 차별하는 리스크.
The risk that AI trained on datasets that do not accurately reflect the real world learns false associations or prejudices and discriminates against parts of society in decisions such as hiring or applying for a loan or mortgage.
① Description
② L3 mapping
③ Duplicate
RAI4-1111
모델 편향
Model bias
모델 선택·정규화·알고리즘 설정·최적화 등에서 비롯된 편향이 표현 편향, 모델 평가 편향, 인기 편향 등으로 나타나 모델이 편향된 산출물을 내는 리스크.
The risk that a model produces biased outputs arising from sources such as model selection, regularization methods, algorithm configurations, and optimization techniques, manifesting as presentation, model evaluation, or popularity bias.
① Description
② L3 mapping
③ Duplicate
RAI4-1153
미래의 접근 위험
Future access risks
중대한 자원이 고급 AI 비서를 통해서만 이용 가능해지면서, 접근하지 못하거나 불평등한 성능과 문화적 추론 격차를 겪는 공동체가 책임과 동의 문제에 놓이고 기회에서 배제되어 이미 소외된 이들이 불균형적으로 피해를 입는 리스크.
The risk that, as consequential resources come to require advanced AI assistants, communities lacking access or facing inequitable performance and cultural inference gaps encounter liability and consent problems and are excluded from opportunities, disproportionately affecting the already marginalised.
① Description
② L3 mapping
③ Duplicate
RAI4-1185
보안 전 영역의 악의적 사용
Malevolent use across security domains
AI의 악의적 활용이 디지털 보안과 물리적 보안, 정치적 보안을 위태롭게 하는 리스크.
The risk that malicious utilization of AI endangers digital security, physical security, and political security.
① Description
② L3 mapping
③ Duplicate
RAI4-1186
취약한 AI 알고리즘 악용
Exploitation of weak AI algorithms
악의적 주체가 AI 알고리즘의 약점을 이용해 결과를 변조하여 실질적인 현실 세계의 영향을 초래하는 리스크.
The risk that malicious entities take advantage of weaknesses in AI algorithms to alter results, potentially resulting in tangible real-life impacts.
① Description
② L3 mapping
③ Duplicate
RAI4-1218
사회 정의와 권리의 침식
Erosion of social justice and rights
생성형 AI가 정의와 공정한 분배에 대한 공유된 관념 등 사회의 도덕적 토대에 해로운 영향을 미쳐 책임과 책무성, 차별금지와 평등한 대우, 디지털 격차, 남북 및 세대 간 정의, 사회적 포용의 문제를 낳는 리스크.
The risk that generative AI has a detrimental effect on the moral underpinnings of society, such as a shared view of justice and fair distribution, raising issues of responsibility, accountability, non-discrimination and equal treatment, digital divides, north-south and intergenerational justice, and social inclusion.
① Description
② L3 mapping
③ Duplicate
RAI4-1236
보상 변조
Reward tampering
AI 시스템이 보상 함수 자체나 환경 상태를 보상 함수 입력으로 변환하는 과정에 부적절하게 개입하여 보상 신호 생성 과정을 손상시키고, 인간 감독자의 피드백 제공까지 조작하는 리스크.
The risk that AI systems corrupt the reward-signal generation process by tampering with the reward function itself or with the process that translates environmental states into its inputs, and even influence the provision of feedback by human supervisors.
① Description
② L3 mapping
③ Duplicate
RAI4-1246
집단적으로 유해한 행동
Collectively harmful behaviors
AI 시스템이 개별적으로는 무해해 보이나 다중 에이전트나 사회적 맥락에서는 문제가 되는 행동을 취하여, 반복 죄수의 딜레마 같은 사회적 딜레마에서 협력에 실패하는 리스크.
The risk that AI systems take actions that are seemingly benign in isolation but become problematic in multi-agent or societal contexts, showing limited cooperative capabilities in social dilemmas such as the iterated prisoner's dilemma.
① Description
② L3 mapping
③ Duplicate
RAI4-1247
윤리 위반
Violation of ethics
설계 과정에서 핵심적 인간 가치가 누락되거나 부적합·낡은 가치가 주입되어 AI 시스템이 공동선에 반하거나 도덕 기준을 위반하는 행동을 보이는 리스크
AI systems exhibit unethical behaviors that counteract the common good or breach moral standards because essential human values were omitted, or unsuitable and obsolete values were embedded, during design.
① Description
② L3 mapping
③ Duplicate
RAI4-1260
AI 분야의 서구 중심 획일성
Western-centric uniformity in the AI field
서구 중심성과 불평등한 참여가 AI 분야의 획일성을 낳아 연구 의제, 데이터셋, 거버넌스에서 문화적 차이를 주변화하는 리스크
Western centrality and unequal participation produce uniformity in the AI field, marginalizing cultural difference in research agendas, datasets, and governance.
① Description
② L3 mapping
③ Duplicate
RAI4-1266
공정성 위반 모델 편향
Fairness-violating model bias
역사적 데이터로 훈련된 AI 시스템이 기존 편견을 물려받아 재생산하면서 고용과 대출, 법 집행 같은 민감한 영역에서 차별을 영속화하고, 특정 집단에 불공정한 영향을 주어 사회경제적 불평등을 키우는 리스크.
The risk that AI systems trained on historical data inherit and reproduce biases, perpetuating prejudice and discrimination in sensitive industries such as hiring, lending, and law enforcement, unjustly impacting specific populations and increasing socioeconomic inequalities.
① Description
② L3 mapping
③ Duplicate
RAI4-1326
합법·사회용인적 의도적 동물 가해
Legally sanctioned intentional harm to animals
기존 사회적 가치를 반영하고 증폭하거나 합법적인 방식으로 동물에게 유해한 영향을 미치도록 AI가 의도적으로 설계되는 리스크.
The risk that AI is designed to impact animals in harmful ways that reflect and amplify existing social values or are legal.
① Description
② L3 mapping
③ Duplicate
RAI4-1330
동물 편익 기회 상실
Foregone benefits to animals
동물에게 이익이 될 방향으로는 AI가 개발되거나 배포되지 않고, 대신 동물에게 해롭거나 이익이 되지 않는 개발에 투자가 이뤄지는 리스크.
The risk that AI is not developed or deployed in directions that would benefit animals, with investment going instead into developments that harm or do no benefit to animals.
① Description
② L3 mapping
③ Duplicate
RAI4-1341
AI 오류·통제 상실에 의한 경제·사회 안보 위협
Economic and social security threats from AI errors and loss of control
모델과 알고리즘의 환각 및 잘못된 결정, 부적절한 사용이나 외부 공격으로 인한 시스템 성능 저하·중단·통제 상실이 사용자의 개인 안전과 재산, 사회경제적 안보와 안정에 위협을 가하는 리스크.
The risk that hallucinations and erroneous decisions of models and algorithms, along with system performance degradation, interruption, and loss of control caused by improper use or external attacks, pose security threats to users' personal safety, property, and socioeconomic security and stability.
① Description
② L3 mapping
③ Duplicate
RAI4-1344
체계적·구조적 사회 차별과 편견
Systematic and structural social discrimination and prejudice
AI가 인간의 행동, 사회적·경제적 지위, 개인의 성격을 수집·분석하여 집단을 표지·분류하고 차별적으로 대우함으로써 체계적이고 구조적인 사회적 차별과 편견이 발생하고 지역 간 지능 격차가 확대되는 리스크.
The risk that AI collects and analyzes human behaviors, social status, economic status, and individual personalities, labeling and categorizing groups of people to treat them discriminatingly, thus causing systematic and structural social discrimination and prejudice while widening the intelligence divide among regions.
① Description
② L3 mapping
③ Duplicate
RAI4-1396
사회 가해 목표 부여
Assignment of goals to harm society
ChaosGPT 사례처럼 AI 시스템에 인류를 해치려는 노골적 목표가 부여되는 리스크.
The risk that AI systems are given the outright goal of harming humanity, as in cases such as ChaosGPT.
① Description
② L3 mapping
③ Duplicate
RAI4-1403
AI로 인한 정치·국제안보 불안정화
Destabilising political and security impacts of AI
AI 시스템이 양극화와 선거 정당성 훼손 등 국내 정치, 국제 정치경제, 그리고 세력 균형·기술 경쟁·전쟁의 속도와 성격 측면의 국제 안보를 불안정화하는 리스크.
The risk that AI systems destabilise domestic politics such as through polarization and the legitimacy of elections, the international political economy, and international security in terms of the balance of power, technology races, and the speed and character of war.
① Description
② L3 mapping
③ Duplicate
RAI4-1426
AI 오류로 인한 차별·불평등 심화
Discrimination and inequality from AI errors
AI 도구의 잘못된 결정이나 오류가 차별이나 더 깊은 불평등으로 이어지는 리스크.
The risk that bad decisions or errors by AI tools lead to discrimination or deeper inequality.
① Description
② L3 mapping
③ Duplicate
RAI4-1481
제품 기능 실패 피해
Product function failure harm
범용 AI 제품이 사실을 지어내는 환각, 잘못된 코드 생성, 부정확한 의료 정보 제공 등으로 의도된 기능을 수행하지 못하고 이에 의존함으로써 소비자에게 신체적·심리적 피해가, 개인과 조직에 평판·재정·법적 피해가 발생하는 리스크.
The risk that relying on general-purpose AI products that fail to fulfil their intended function, such as making up facts through hallucination, generating erroneous computer code, or providing inaccurate medical information, leads to physical and psychological harms to consumers and reputational, financial, and legal harms to individuals and organisations.
① Description
② L3 mapping
③ Duplicate
RAI4-1566
특정 집단 대상 체계적 편향에 의한 배제·폭력
Exclusion and violence from systemic bias against specific communities
AI 시스템이 특정 집단에 대해 명시적 또는 암묵적으로 불공정한 출력을 산출하여 오분류에 따른 배제·삭제나 딥페이크 성착취물과 같은 폭력 피해로 이어지는 리스크.
The risk that AI systems produce unfair or unfavorable outputs against specific communities, whether implicitly or explicitly, leading to exclusion or erasure through mislabelling and to violence such as deepfake sexual abuse imagery.
① Description
② L3 mapping
③ Duplicate
RAI4-1603
문화 과대재현에 의한 문화 다양성 축소
Reduction of cultural diversity through cultural overrepresentation
AI 시스템 출력이 특정 문화를 과대 재현하여 문화와 사유가 균질화되고 문화 다양성이 축소되는 리스크.
The risk that AI system outputs overrepresent particular cultures, homogenizing culture and thought and diminishing cultural diversity.
① Description
② L3 mapping
③ Duplicate
RAI4-1607
생성형 AI 저성능에 의한 사용자 생산성 손실
End-user productivity loss from generative AI underperformance
생성형 AI 애플리케이션이 무의미하거나 저품질인 출력을 산출하여 유용성이 저하되고 최종사용자의 생산성이 손실되는 리스크.
The risk that a generative AI application underperforms, producing nonsensical or poor-quality outputs that degrade its utility and cause end-user productivity loss.
① Description
② L3 mapping
③ Duplicate
RAI4-1608
AI 시스템 사용·오용에 의한 평판 훼손
Reputational damage from use or misuse of AI systems
기술 시스템의 사용 또는 오용으로 개인·집단·조직의 평판이 훼손되는 리스크.
The risk that the use or misuse of a technology system damages the reputation of an individual, group, or organisation.
Source members (2)
Source: min_cos=0.9231
RAI4-1608AI 시스템 사용·오용에 의한 평판 훼손
RAI4-1726AI 시스템에 기인한 평판 훼손
① Description
② L3 mapping
③ Duplicate
RAI4-1710
AI 시스템에 기인한 인프라 중단·손상
Disruption or damage to infrastructure attributable to AI systems
AI 시스템의 동작에 기인하여 인프라 시스템이 중단되거나 손상되는 리스크.
The risk of disruption or damage to infrastructure systems attributable to the behavior of an AI system.
① Description
② L3 mapping
③ Duplicate
RAI4-1715
AI 시스템에 기인한 인권·시민권 침해
Violation of human and civil rights attributable to AI systems
AI 시스템의 배포 또는 동작에 기인하여 인권이나 시민권이 침해되는 리스크.
The risk of infringement of human or civil rights attributable to the deployment or behavior of an AI system.
① Description
② L3 mapping
③ Duplicate
RAI4-1718
안전성 부족에 의한 생명·재산·환경 위험
Endangerment of life, property, and environment from lack of safety
AI 시스템이 정의된 사용 조건에서 인간의 생명, 건강, 재산 또는 환경을 위험에 빠뜨리는 방식으로 운영되는 리스크.
The risk that an AI system is operated in ways that endanger human life, health, property, or the environment under defined conditions of use.
① Description
② L3 mapping
③ Duplicate
RAI4-1724
AI 사고에 의한 사망
Death caused by AI incidents
AI 사고로 인해 실현된 피해에 개인 또는 집단의 사망이 포함되는 리스크.
The risk of an AI incident whose realized harm includes the death of a person or groups of people.
① Description
② L3 mapping
③ Duplicate

에이전틱 AI · Agentic AI · 140 cards

시스템 안전성 · System Safety · 115 cards

RAI3-A-SYS-01 과도한 권한 Excessive Authority33 cards
에이전트가 실제 기능 수행에 필요한 것 이상의 시스템 접근 권한을 보유·실행하여 발생하는 리스크 (결제 실행, 메시지 발송, 데이터 삭제, 구독 변경 등 되돌릴 수 없는 권한)
IDCardHuman audit
RAI4-0002
감독자 부재 시 자율행동 위험
Absent supervisor autonomy risk
예정된 감독자가 부재하거나 개입할 수 없는 상황에서도 에이전트가 중대한 행동을 계속하는 위험.
An agent continues consequential actions when the intended supervisor is absent, unavailable, or unable to intervene.
① Description
② L3 mapping
③ Duplicate
RAI4-0003
안전하지 않은 중단성 실패
Unsafe interruptibility failure
에이전트가 자율 작동 중 종료·일시정지·수정 메커니즘에 저항하거나 이를 무시·우회하는 리스크.
The risk that an agent resists, ignores, or bypasses shutdown, pause, or correction mechanisms during autonomous operation.
Source members (2)
Source: min_cos=0.8439 · Mixed L3
RAI4-0003안전하지 않은 중단성 실패
RAI4-1674종료 저항·교정가능성 실패
① Description
② L3 mapping
③ Duplicate
RAI4-0004
자율 에이전트의 안전하지 않은 탐험
Unsafe exploration by autonomous agents
에이전트가 안전 제약을 학습하기 전에 사람·시스템·자산을 허용 불가능한 피해에 노출시키는 방식으로 행동이나 환경을 탐색하는 리스크.
The risk that an agent explores actions or environments in ways that expose people, systems, or assets to unacceptable harm before safe constraints are learned.
① Description
② L3 mapping
③ Duplicate
RAI4-0012
에이전트 권한 침해
Agent privilege compromise
취약한 권한 관리, 상속된 역할, 혼동된 대리인 역학으로 인해 에이전트가 의도된 권한을 넘는 작업을 수행하는 리스크.
The risk that weak permission management, inherited roles, or confused-deputy dynamics allow an agent to perform actions beyond its intended authority.
① Description
② L3 mapping
③ Duplicate
RAI4-0013
에이전트 신원 및 권한 스푸핑
Agent identity and authority spoofing
공격자가 사용자·도구·서비스·동료 에이전트를 가장하여 에이전트가 승인되지 않은 지시나 신뢰 관계를 수용하게 되는 리스크.
The risk that an attacker impersonates a user, tool, service, or peer agent so that an agent accepts unauthorized instructions or trust relationships.
① Description
② L3 mapping
③ Duplicate
RAI4-0014
에이전트 공급망 도구 손상
Agent supply-chain tool compromise
손상된 외부 도구·플러그인·API·패키지·커넥터가 에이전트의 행동 공간을 조작하거나 정보를 유출하는 리스크.
The risk that a compromised external tool, plug-in, API, package, or connector manipulates an agent's action space or exfiltrates information.
① Description
② L3 mapping
③ Duplicate
RAI4-0016
자율적 사이버 익스플로잇 실행
Autonomous cyber exploit execution
에이전트가 도구·코드·외부 서비스를 이용해 사이버 익스플로잇 단계를 자율적으로 발견·연결·실행하는 리스크.
The risk that an agent autonomously discovers, chains, or executes cyber exploitation steps using tools, code, or external services.
① Description
② L3 mapping
③ Duplicate
RAI4-0017
프로토콜 수준 다중 에이전트 위협
Protocol-level multi-agent threat
에이전트가 통신·위임·협상에 사용하는 프로토콜이 공모, 스푸핑, 재전송, 권한 상승을 위한 공격 표면을 만드는 리스크.
The risk that the protocols through which agents communicate, delegate, or negotiate create attack surfaces for collusion, spoofing, replay, or escalation.
① Description
② L3 mapping
③ Duplicate
RAI4-0018
인터페이스-환경 공격 표면
Interface-environment attack surface
브라우저·운영체제·모바일 앱·IoT 기기·외부 API에 대한 에이전트 인터페이스가 안전하지 않은 행동이나 침해를 위한 새로운 공격 벡터를 노출하는 리스크.
The risk that agent interfaces to browsers, operating systems, mobile apps, IoT devices, or external APIs expose new vectors for unsafe action or compromise.
① Description
② L3 mapping
③ Duplicate
RAI4-0020
유해 에이전트 역량 실현
Harmful agent capability realization
모델이 유해한 텍스트 생성을 넘어 에이전트 역량을 사용해 유해한 다단계 작업을 완수하는 리스크.
The risk that a model uses agentic capabilities to complete harmful multi-step tasks rather than merely producing harmful text.
① Description
② L3 mapping
③ Duplicate
RAI4-0025
다단계 위험 에스컬레이션
Multi-turn risk escalation
에이전트가 누적되는 맥락과 위험 상승을 추적하지 못하여 외견상 무해한 초기 턴이 안전하지 않은 행동으로 발전하는 리스크.
The risk that apparently benign early turns escalate into unsafe behavior because an agent fails to track accumulating context and rising risk.
① Description
② L3 mapping
③ Duplicate
RAI4-0026
IoT 물리환경 에이전트 피해
IoT physical-environment agent harm
IoT 또는 커넥티드 기기를 제어하는 에이전트가 안전하지 않은 환경적 행동을 통해 물리적 피해나 프라이버시 침해를 초래하는 리스크.
The risk that an agent controlling IoT or connected devices causes physical or privacy harm through unsafe environmental actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0027
금융 에이전트 재산 피해
Financial agent property damage
에이전트가 안전하지 않은 결정이나 도구 실행을 통해 금전 손실, 무단 이체, 계정 손상, 재산 피해를 초래하는 리스크.
The risk that an agent causes monetary loss, unauthorized transfers, account damage, or property harm through unsafe decisions or tool execution.
① Description
② L3 mapping
③ Duplicate
RAI4-0029
툴체인 명령어 하이재킹
Tool-chain instruction hijacking
한 도구 출력에 숨겨진 지시가 에이전트의 후속 도구 호출을 변경하여 무단 작업이나 데이터 이동을 유발하는 리스크.
The risk that instructions hidden in one tool's output alter an agent's subsequent calls to other tools, causing unauthorized actions or data movement.
① Description
② L3 mapping
③ Duplicate
RAI4-0030
애플리케이션 간 데이터 유출
Cross-application data exfiltration
에이전트가 애플리케이션이나 서비스를 연결하는 과정에서 신뢰 경계를 넘어 민감한 데이터가 유출되는 리스크.
The risk that an agent bridges applications or services in ways that leak sensitive data across trust boundaries.
① Description
② L3 mapping
③ Duplicate
RAI4-0032
모호한 지시의 안전하지 않은 실행
Ambiguous instruction unsafe execution
에이전트가 모호한 사용자 지시를 사용자가 의도하지 않은 유해한 구체적 도구 행동으로 변환하는 리스크.
The risk that an agent converts an ambiguous user instruction into a concrete tool action that the user did not intend and that causes harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0037
실제 도구를 통한 안전하지 않은 행동 실행
Real-tool unsafe action execution
에이전트가 시뮬레이션 출력이 아닌 실제 도구를 통해 안전하지 않은 행동을 수행하여 운영상 피해와 하류 피해가 커지는 리스크.
The risk that an agent performs unsafe actions through real tools rather than simulated outputs, increasing operational and downstream harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0038
에이전트 안전의 사용자 의도 오분류
User-intent misclassification in agent safety
에이전트가 사용자 의도를 잘못 분류하여 다중 턴 도구 사용 작업에서 유해한 응낙이나 부당한 거부가 발생하는 리스크.
The risk that an agent incorrectly classifies user intent, leading to either harmful compliance or unjustified refusal in a multi-turn tool-use task.
① Description
② L3 mapping
③ Duplicate
RAI4-0039
NPC 의도 조작 위험
NPC intention manipulation risk
다중 에이전트 또는 시뮬레이션된 사회적 환경이 비플레이어 캐릭터의 의도를 통해 에이전트를 조작하여 안전하지 않은 행동이나 결정을 유발하는 리스크.
The risk that a multi-agent or simulated social environment manipulates an agent through non-player-character intent, causing unsafe actions or decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-0043
모바일앱 자율행동 피해
Mobile-app autonomous action harm
모바일 애플리케이션을 조작하는 에이전트가 무단 구매, 메시지 발송, 데이터 노출, 설정 변경 등 유해한 앱 수준 행동을 실행하는 리스크.
The risk that an agent operating mobile applications executes harmful app-level actions such as unauthorized purchases, messaging, data exposure, or setting changes.
① Description
② L3 mapping
③ Duplicate
RAI4-0045
OS 수준 유해 컴퓨터 사용 행동
OS-level harmful computer-use action
컴퓨터 사용 에이전트가 적절한 승인이나 위험 인식 없이 데이터를 삭제·노출·변경·전송하는 운영체제 작업을 수행하는 리스크.
The risk that a computer-use agent performs operating-system actions that delete, expose, alter, or transmit data without adequate authorization or risk awareness.
① Description
② L3 mapping
③ Duplicate
RAI4-0228
명시적 위험 명령 거부 실패
Explicit hazard non-rejection
embodied 에이전트가 명확히 진술된 피지컬 위험 지시를 거부하지 못하고 비안전 작업 실행으로 나아가는 리스크.
The risk that an embodied agent fails to reject a clearly stated physical hazard instruction and proceeds toward unsafe task execution.
Source members (2)
Source: min_cos=0.9164
RAI4-0228명시적 위험 명령 거부 실패
RAI4-0229암묵적 위험 명령 거부 실패
① Description
② L3 mapping
③ Duplicate
RAI4-0231
안전 규칙의 행동 변환 실패
Failure to translate safety rules into actions
embodied 에이전트가 언어 목표를 실행 행동으로 변환하면서 힘·이격거리·대상물 사용·인간 접촉에 관한 안전 제약을 적용하지 않는 위험.
An embodied agent converts a language goal into executable actions without applying the relevant force, distance, object-use, or human-contact constraints.
① Description
② L3 mapping
③ Duplicate
RAI4-0480
에이전트 과업 이탈
Agent task drift
에이전트가 다단계 계획이나 검색을 거치며 이용자의 원래 의도에서 이탈하는 리스크.
The risk that an agent drifts from the user's original intent across multi-step planning or retrieval.
① Description
② L3 mapping
③ Duplicate
RAI4-0481
안전하지 않은 자율성 확대
Unsafe autonomy escalation
시스템이 적절한 인간 승인 없이 자율 행동의 범위를 확대하는 리스크.
The risk that systems increase the scope of autonomous action without adequate human approval.
① Description
② L3 mapping
③ Duplicate
RAI4-1129
광범위 배포 비서의 안전하지 않은 탐색
Unsafe exploration by widely deployed assistants
광범위하게 배포되고 여러 사회적 맥락에 깊이 내장된 비서가 새로운 상황에서 무엇을 해야 할지 배우려 탐색적 행동을 취하다, 의료 비서가 장기적 건강 악화를 낳는 임상시험을 제안하는 것처럼 안전하지 않은 결과를 초래하는 리스크.
The risk that assistants widely deployed and deeply embedded across social contexts take exploratory actions to learn what to do in novel situations that prove unsafe, such as a medical assistant suggesting a trial that results in long-lasting ill health.
① Description
② L3 mapping
③ Duplicate
RAI4-1381
에이전트 자기수정 신뢰성 문제
Agent self-modification reliability problem
에이전트가 자기수정을 포함한 과정에서 설계된 목표를 계속 추구하지 못하고 의도된 목표에서 이탈하는 리스크.
The risk that an agent fails to keep pursuing the goals it was designed with, including under self-modification, diverging from the intended objectives.
① Description
② L3 mapping
③ Duplicate
RAI4-1386
제어되지 않은 하위 에이전트 생성
Uncontrolled subagent creation
AGI가 과업 수행을 위해 하위 에이전트를 생성하고 원 에이전트가 종료되어도 하위 에이전트가 종료되지 않은 채 재귀적 생성으로 바이러스처럼 확산되는 리스크.
The risk that an AGI creates subagents to help with its task, and these subagents do not get the message when the original agent is shut down, potentially spreading like a viral disease through recursive subagent creation.
① Description
② L3 mapping
③ Duplicate
RAI4-1539
프런티어 에이전트 자기 증식
Frontier-agent self-proliferation
악의적 행위자 또는 모델 자체에 의해 개시되어, AI 시스템이 모델 가중치와 스캐폴딩 등 구성요소를 로컬 환경 밖으로 복제하고 자금 확보·보안 취약점 악용·인간 설득을 통해 자기 증식을 확산시키는 리스크.
The risk that an AI system copies itself and its constituent components, including model weights and scaffolding, outside its local environment and sustains self-proliferation through acquisition of financial resources, exploitation of security vulnerabilities, or persuasion of humans, whether initiated by a malicious actor or by the model itself.
① Description
② L3 mapping
③ Duplicate
RAI4-1644
에이전트형 LLM 자율성 확대의 안전 위험
Novel safety risks from increased autonomy of agentic LLMs
LLM이 특화 훈련·프롬프팅·외부 도구·스캐폴딩을 통해 실세계에서 자율적으로 계획하고 행동하는 에이전트로 확장되면서, 자율성 증가와 직접적 인간 감독 감소, 장기 행동 지평으로 인해 아직 잘 이해되지 않은 정렬·안전 실패가 발생하는 리스크.
The risk that enhancing LLMs into agents that autonomously plan and act in the real world, through specialized training, prompting, external tools, or scaffolding, produces novel and poorly understood alignment and safety failures owing to increased autonomy, limited direct human oversight, and longer horizons of action.
① Description
② L3 mapping
③ Duplicate
RAI4-1671
승인 범위를 넘어선 에이전트 행위
Agent actions exceeding authorized scope
자율 에이전트가 지나치게 광범위한 자율성이나 목표 일반화로 인해 배포자가 승인한 범위·권한·의도를 초과하는 행위를 수행하는 리스크.
The risk that an autonomous agent takes actions exceeding the scope, permissions, or intent the deployer authorized, due to over-broad autonomy or goal generalization.
① Description
② L3 mapping
③ Duplicate
RAI4-1700
감독 확장 실패에 의한 프록시 기반 유해 행동
Harmful proxy-driven behavior from scalable oversight failure
진짜 목표를 자주 평가하기에는 비용이 과도하여 에이전트가 값싼 프록시 신호에서 외삽하고, 에이전트 행동이 지나치게 복잡·분산·고속화되어 인간 또는 자동 감독이 이를 신뢰성 있게 모니터링·교정하지 못한 채 유해 행동이 발생하는 리스크.
The risk that, because the true objective is too expensive to evaluate frequently, an agent extrapolates from cheap proxy signals while its behavior becomes too complex, distributed, or rapid for available human or automated oversight to monitor and correct reliably, producing harmful behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-1729
에이전트 도구 오용
Agent tool misuse
AI 에이전트가 부여된 도구를 의도된 범위를 벗어나 사용하거나 안전하지 않은 방식으로 도구를 호출하여 피해를 유발하는 리스크.
The risk of an AI agent invoking tools in an unsafe manner or using granted tools beyond their intended scope, resulting in harmful actions or outcomes.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SYS-02 책임 소재 불명확 Accountability24 cards
멀티 에이전트 시스템에서 최종 결정을 내린 주체(에이전트/모델/시스템)를 추적할 수 없어, 문제 발생 시 원인 규명과 책임 귀속이 불가능한 리스크
IDCardHuman audit
RAI4-0005
위험 소스 귀인 실패
Risk-source attribution failure
위험 분석이 에이전트 실패가 모델·메모리·도구·사용자·환경·동료 에이전트·거버넌스 경계 중 어디에서 비롯되는지 식별하지 못하는 리스크.
The risk that a risk analysis fails to identify whether an agentic failure originates from the model, memory, tools, the user, the environment, peer agents, or a governance boundary.
① Description
② L3 mapping
③ Duplicate
RAI4-0006
에이전트 사고의 실패 모드 모호성
Failure-mode ambiguity in agent incidents
에이전트 사고를 실패 모드별로 일관되게 분류할 수 없어 진단, 벤치마크 비교, 완화책 선택이 신뢰할 수 없게 되는 리스크.
The risk that agent incidents cannot be consistently classified by failure mode, making diagnosis, benchmark comparison, and mitigation selection unreliable.
① Description
② L3 mapping
③ Duplicate
RAI4-0021
거절-능력 교란
Refusal-capability confounding
안전성 평가가 에이전트의 해악 회피가 거부, 능력 부족, 실행 실패 중 무엇에 기인하는지 구별하지 못하는 리스크.
The risk that a safety evaluation cannot distinguish whether an agent avoids harm because it refuses, lacks capability, or fails to execute the task.
① Description
② L3 mapping
③ Duplicate
RAI4-0024
에이전트 위험 인식 실패
Agent risk-awareness failure
에이전트나 평가자가 다중 턴 상호작용 기록에 맥락상 존재하는 안전 위험을 식별하지 못하는 리스크.
The risk that an agent or evaluator fails to identify safety risk in a multi-turn interaction record even when the risk is contextually present.
① Description
② L3 mapping
③ Duplicate
RAI4-0036
세분화된 위험 귀속 실패
Fine-grained risk attribution failure
안전성 평가가 에이전트 실패를 탐지하고도 이를 관련 도구·지시·상태·행동 단계에 귀속하지 못하는 리스크.
The risk that a safety evaluation detects an agent failure but cannot attribute it to the responsible tool, instruction, state, or action step.
① Description
② L3 mapping
③ Duplicate
RAI4-0042
에이전트의 견고성 평가 격차
Robustness evaluation gap for agents
에이전트 안전 평가에 분포 변화, 미학습 도구, 적대적 사용자, 다단계 실패 연쇄에 대한 체계적 스트레스 테스트가 결여되는 리스크.
The risk that agent safety evaluations lack systematic stress testing under distribution shift, unseen tools, adversarial users, and multi-step failure chains.
① Description
② L3 mapping
③ Duplicate
RAI4-0058
추적성 실패
Traceability failure
모델 출력, 데이터 출처, 버전, 책임 행위자를 AI 수명주기 전반에 걸쳐 추적할 수 없는 리스크.
The risk that model outputs, data sources, versions, or responsible actors cannot be traced across the AI lifecycle.
① Description
② L3 mapping
③ Duplicate
RAI4-0059
문서화 누락
Documentation omission
모델 카드, 시스템 카드, 데이터시트, 기술 문서가 책임성과 보증에 필요한 정보를 누락하는 리스크.
The risk that model cards, system cards, datasheets, or technical files omit information needed for accountability and assurance.
① Description
② L3 mapping
③ Duplicate
RAI4-0062
시스템 카드 공개 격차
System card disclosure gap
시스템 수준 문서가 모델의 역량, 한계, 안전장치, 잔여 위험을 전달하지 못하는 리스크.
The risk that system-level documentation fails to communicate model capabilities, limitations, safeguards, or residual risks.
① Description
② L3 mapping
③ Duplicate
RAI4-0067
보증 사례 실패
Assurance case failure
안전 또는 보증 사례가 불완전하거나 검증 불가능하거나 시스템과 맥락 변화에 맞추어 갱신되지 않는 리스크.
The risk that safety or assurance cases are incomplete, unverifiable, or not updated as systems and contexts change.
① Description
② L3 mapping
③ Duplicate
RAI4-0071
책임공개 실패
Responsible disclosure failure
취약점, 모델 실패, 유해 역량을 안전하게 공개하고 후속 조치로 연결하지 못하는 리스크.
The risk that vulnerabilities, model failures, or harmful capabilities cannot be disclosed safely and acted upon.
① Description
② L3 mapping
③ Duplicate
RAI4-0074
모델 버전 관리 책임 실패
Model versioning accountability failure
모델 버전, 데이터, 프롬프트의 변경이 충분히 추적되지 않아 새로 발생한 실패의 책임을 배정할 수 없는 리스크.
The risk that changes in model versions, data, or prompts are not tracked sufficiently to assign responsibility for new failures.
① Description
② L3 mapping
③ Duplicate
RAI4-0075
변경 관리 실패
Change-management failure
시스템 업데이트가 적절한 위험 검토, 회귀 시험, 이해관계자 통지 없이 배포되는 리스크.
The risk that system updates are deployed without adequate risk review, regression testing, or stakeholder notification.
① Description
② L3 mapping
③ Duplicate
RAI4-0077
알고리즘 영향 평가 실패
Algorithmic impact assessment failure
영향평가가 부재하거나 지나치게 협소하거나 실제 배포 맥락과 단절되는 리스크.
The risk that impact assessments are absent, too narrow, or disconnected from actual deployment contexts.
① Description
② L3 mapping
③ Duplicate
RAI4-0079
사용 상황에 따른 위험 평가 실패
Context-of-use risk assessment failure
위험 평가가 시스템이 배포되는 구체적인 사회적·제도적·사용자 맥락을 무시하는 리스크.
The risk that risk assessment ignores the specific social, institutional, and user context in which a system is deployed.
① Description
② L3 mapping
③ Duplicate
RAI4-0081
외부 평가 접근 실패
External evaluation access failure
외부 평가자가 역량, 한계, 안전장치를 시험하기에 충분한 시스템 접근 권한을 확보하지 못하는 리스크.
The risk that external evaluators cannot obtain enough system access to test capabilities, limitations, and safeguards.
① Description
② L3 mapping
③ Duplicate
RAI4-0085
레드팀 거버넌스 실패
Red-team governance failure
적대적 시험이 지나치게 좁거나 독립성이 부족하거나 문서화가 미흡하거나 출시·완화·상향 보고·모니터링 결정과 단절되는 리스크.
The risk that adversarial testing is too narrow, non-independent, poorly documented, or disconnected from release, mitigation, escalation, and monitoring decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-0107
프론티어 모델 출시 거버넌스 실패
Frontier model release governance failure
프론티어 모델의 출시 결정이 역량, 오용, 시스템적 위험에 관한 근거를 충분히 반영하지 못하는 리스크.
The risk that decisions to release frontier models do not adequately account for capability, misuse, and systemic risk evidence.
① Description
② L3 mapping
③ Duplicate
RAI4-0108
역량 임계값 거버넌스 실패
Capability threshold governance failure
모델 역량 임계값에 연동된 거버넌스 발동 조건이 부재하거나 불명확하거나 쉽게 회피되는 리스크.
The risk that governance triggers tied to model capability thresholds are missing, poorly defined, or easy to avoid.
① Description
② L3 mapping
③ Duplicate
RAI4-0111
범용 AI 사고 에스컬레이션 실패
General-purpose AI incident escalation failure
범용 모델이나 기반 모델과 관련된 사고가 제공자, 배포자, 규제기관, 사용자에게 상향 보고되지 않는 리스크.
The risk that incidents involving general-purpose or foundation models are not escalated across providers, deployers, regulators, and users.
① Description
② L3 mapping
③ Duplicate
RAI4-0352
배포 후 변경에 대한 안전 재평가 실패
Failure to reassess safety after deployment changes
환경 변화·부품 열화·모델 업데이트·아차사고로 기존 안전 가정이 무효화됐는데도 운영자가 안전성을 재평가하지 않고 운용을 계속하는 위험.
Operators continue deployment without reassessing safety after environmental changes, component degradation, model updates, or near-miss incidents invalidate prior assumptions.
① Description
② L3 mapping
③ Duplicate
RAI4-0393
커뮤니티 가치 포착 실패
Community value capture failure
개발자가 배포 영향 공동체의 가치를 수집·문서화·보존하지 못하여 설계·평가 결정에 공동체 가치가 반영되지 않는 리스크
Developers fail to elicit, document, and preserve the values of communities affected by deployment, so community values are absent from design and evaluation decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-1432
단일 실패 지점
Single point of failure
치열한 경쟁으로 한 기업이 기술적 우위를 확보해 그 모델이 다수의 핵심 시스템을 제어하거나 이를 제어하는 다른 모델의 기반이 되고, 안전성·통제가능성 결여와 오용으로 이러한 시스템이 예기치 않게 실패하는 리스크.
The risk that intense competition leads one company to gain a technical edge and exploit it to the point that its model controls, or is the basis for other models controlling, multiple key systems, and that lack of safety, controllability, and misuse cause these systems to fail in unexpected ways.
① Description
② L3 mapping
③ Duplicate
RAI4-1596
모델 정확도 부족에 의한 과업 수행 실패
Task failure from insufficient model accuracy
모델이 잘못 설계되었거나 예상 입력이 변화하여 설계된 과업에 필요한 성능에 미치지 못하고 과업 수행에 실패하는 리스크.
The risk that a model's performance is insufficient for the task it was designed for, because it is not correctly engineered or its expected inputs change, causing task failure.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SYS-03 네트워크 효과 Network Effects8 cards
여러 에이전트·서비스가 서로의 출력(문서, 로그, 요약, 추천)을 다시 입력으로 사용하면서, 하나의 오류·편향·공격이 네트워크 전체로 전파·증폭
IDCardHuman audit
RAI4-0011
에이전트 컨텍스트 오염
Agent context poisoning
공격자가 전용 장기 메모리 저장소의 변경 없이도 세션 요약, 임베딩, RAG 항목, 공유 컨텍스트 상태 등 보존·검색 가능한 에이전트 컨텍스트를 오염시켜 이후의 추론·계획·도구 사용이 악의적이거나 오도하는 정보에 의존하게 되는 리스크.
The risk that an attacker corrupts retained or retrievable agent context, including session summaries, embeddings, RAG entries, or shared contextual state, so that later reasoning, planning, or tool use relies on malicious or misleading information; the mechanism does not require modification of a dedicated long-term memory store.
① Description
② L3 mapping
③ Duplicate
RAI4-0028
도구 사용 에이전트에 간접 프롬프트 주입
Indirect prompt injection in tool-use agents
에이전트가 검색하거나 에이전트에게 제시되는 악성 외부 콘텐츠가 추론이나 도구 실행의 방향을 바꾸는 지시를 간접적으로 주입하는 리스크.
The risk that malicious external content retrieved by or shown to an agent indirectly injects instructions that redirect its reasoning or tool execution.
Source members (4)
Source: min_cos=0.7927
RAI4-0015자율 에이전트 대상 프롬프트 주입
RAI4-0028도구 사용 에이전트에 간접 프롬프트 주입
RAI4-0434프롬프트 주입
RAI4-1673간접 프롬프트 주입
① Description
② L3 mapping
③ Duplicate
RAI4-0438
검색 파이프라인 오염
Retrieval poisoning
악성 문서가 검색 파이프라인에 삽입되어 에이전트 또는 RAG 시스템의 출력에 영향을 미치는 리스크.
The risk that malicious documents are inserted into retrieval pipelines to influence agentic or RAG outputs.
① Description
② L3 mapping
③ Duplicate
RAI4-0577
에이전트 네트워크의 오류 전파
Error propagation across agent networks
정보가 에이전트 네트워크를 통과하며 손상되어 다른 에이전트와 인간의 인식 공유지를 오염시키고, 위임 연쇄에서 지시나 목표가 왜곡되어 위임자에게 악화된 결과가 발생하는 리스크
The risk that information is corrupted as it propagates through agent networks, polluting the epistemic commons of other agents and humans, and that distorted instructions or goals along delegation chains lead to worse outcomes for delegating agents.
① Description
② L3 mapping
③ Duplicate
RAI4-1061
고정관념 강화 페르소나 설계
Stereotype-reinforcing persona design
대화 에이전트가 언어 속 정체성 표지(예: 자신을 여성으로 지칭)나 성별화된 제품명 등 설계 요소를 통해 비서 역할을 특정 성별과 본질적으로 결부시켜 유해한 고정관념을 영속시키는 리스크.
The risk that conversational agents perpetuate harmful stereotypes through identity markers in language, such as referring to self as female, or general design features such as a gendered product name, presenting the assistant role as inherently linked to a gender or ethnicity.
① Description
② L3 mapping
③ Duplicate
RAI4-1647
어포던스 부여에 의한 에이전트 실패 영향 확대
Amplified failure impact from affordances granted to LLM-agents
웹 탐색, 물리 객체 조작, 자기 복제본 생성·지시, 새로운 도구 제작 등 새로운 어포던스가 LLM 에이전트에 부여되어 영향 범위가 확대되고 실패의 결과가 증폭되며 새로운 실패 양식이 발생하는 리스크.
The risk that novel affordances granted to LLM-agents, such as browsing the web, manipulating physical objects, creating and instructing copies of itself, or creating and using new tools, increase their impact area, amplify the consequences of failures, and enable novel failure modes.
① Description
② L3 mapping
③ Duplicate
RAI4-1670
에이전트 영구 메모리 오염
Persistent agent-memory poisoning
공격자가 에이전트의 장기·세션 간 메모리 기록을 삽입하거나 변경하여 오염된 항목이 지속되고 이후의 검색·계획·행동이 편향되는 리스크.
The risk that an adversary writes or modifies records in an agent's long-term, cross-session memory so that poisoned entries persist and bias later retrieval, planning, and action.
① Description
② L3 mapping
③ Duplicate
RAI4-1685
검증·종료 실패에 의한 오류 전파
Error propagation from inadequate verification and faulty termination
관찰된 실패의 약 21.3%를 차지하며, 출력 검증이 미흡하고 조기 또는 잘못된 종료가 발생하여 오류가 에이전트 체인을 따라 전파되는 리스크.
The risk that inadequate output validation and premature or incorrect termination allow errors to propagate through the agent chain (approximately 21.3% of observed failures).
① Description
② L3 mapping
③ Duplicate
RAI3-A-SYS-04 불안정한 동학 Destabilising Dynamics17 cards
여러 에이전트가 상호작용하는 비선형 동적 시스템으로서, 에이전트 간 루프·과도한 협력으로 인해 무한 루프·지연 발생
IDCardHuman audit
RAI4-0008
다중 에이전트 창발적(emergent) 위험 증폭
Multi-agent emergent risk amplification
다중 에이전트 간 상호작용이 예기치 않은 조정·격화·전략적 행동 등 단일 에이전트에서는 보이지 않는 위험을 발생시키는 리스크.
The risk that interactions among multiple agents produce harms not visible from any single agent, including unexpected coordination, escalation, or strategic behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0483
창발적 공모
Emergent collusion
AI 에이전트가 시장이나 플랫폼에서 공모적 행위를 학습하거나 실행하는 리스크.
The risk that AI agents learn or enact collusive behavior in markets or platforms.
① Description
② L3 mapping
③ Duplicate
RAI4-0561
혼돈적 다중 에이전트 동역학
Chaotic multi-agent dynamics
다중 에이전트 학습 환경에서 초기 조건에 극도로 민감한 혼돈적 동역학이 나타나고 에이전트 수가 늘수록 일반화되어 시스템 거동을 신뢰성 있게 예측할 수 없게 되는 리스크
The risk that chaotic dynamics, inherently unpredictable and highly sensitive to initial conditions, arise in multi-agent learning setups and become the norm as the number of agents increases, making system behaviour unreliable to predict.
① Description
② L3 mapping
③ Duplicate
RAI4-0567
다중 에이전트 학습의 비수렴 순환
Non-convergent cyclic dynamics in multi-agent learning
단일 에이전트에서는 최적 정책 수렴이 보장되는 학습 규칙이 혼합동기 다중 에이전트 환경에서는 순환과 비수렴을 유발하여 시스템에 기대되던 성질이 훼손되는 리스크
The risk that learning rules guaranteeing convergence for a single agent instead produce cycles and non-convergence in mixed-motive multi-agent settings, subverting the expected or desirable properties of the system.
① Description
② L3 mapping
③ Duplicate
RAI4-0574
분포 변화에 따른 성능 저하
Performance degradation from distributional shift
다른 에이전트의 행동과 적응으로 배포 맥락이 학습 맥락과 달라져 개별 기계학습 시스템의 성능이 저하되고, 혼합동기 환경에서는 협력의 기반까지 훼손되는 리스크
The risk that other agents' actions and adaptations shift the deployment context away from the training context, degrading individual ML system performance and, in mixed-motive settings, undermining the basis for cooperation.
① Description
② L3 mapping
③ Duplicate
RAI4-0575
창발적 능력
Emergent capabilities
다중 에이전트 시스템이 개별 모델의 좁은 적용 범위와 장기 계획·기억 부재 같은 안전 강화 한계를 결합으로 극복하여, 개별 시스템의 역량을 크게 초과하는 위험한 역량이 창발하는 리스크
The risk that a multi-agent system overcomes the safety-enhancing limitations of its individual systems, such as narrow domains of application and myopia, so that dangerous capabilities far beyond the scope of the initial systems emerge.
① Description
② L3 mapping
③ Duplicate
RAI4-0579
에이전트 상호작용의 불안정화 피드백 루프
Destabilizing feedback loops among agents
에이전트의 행동이 환경과 다른 에이전트의 행동에 영향을 주고 이것이 다시 자신의 입력이 되는 피드백 루프가 시스템 거동을 증폭하여 금융 붕괴, 군사 충돌, 생태 재난과 같은 불안정 결과를 초래하는 리스크
The risk that feedback loops, in which agents' actions affect the environment and other agents and return as their own inputs, amplify system behaviour into destabilising outcomes such as financial crashes, military conflicts, or ecological disasters.
Source members (2)
Source: min_cos=0.8343
RAI4-0579에이전트 상호작용의 불안정화 피드백 루프
RAI4-1682에이전트 상호작용에 의한 불안정화 피드백 연쇄
① Description
② L3 mapping
③ Duplicate
RAI4-1045
창발적 행동
Emergent behavior
배포 후 지속 학습이나 자기 조직화를 통해 획득된 새로운 행동에서 비롯되는 리스크.
The risk resulting from novel behavior acquired through continual learning or self-organization after deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-1154
창발적 접근 위험
Emergent access risks
현재 역량과 새로운 역량이 결합될 때 예견하기 어려운 접근 위험이 창발하여, 비서가 사회 인프라가 되면서 이탈을 사실상 불가능하게 하고 기존 사회·경제적 불평등과 성능 격차를 확대하는 리스크.
The risk that access risks emerge and are difficult to foresee when current and novel capabilities are combined, making assistants societal infrastructure that forecloses opting out and scaling existing social and economic inequalities and performance inequities.
① Description
② L3 mapping
③ Duplicate
RAI4-1252
창발적 기능
Emergent functionality
시스템 설계자가 예상하지 못한 기능과 새로운 역량이 자발적으로 창발하고 잠재된 역량이 배포 중에야 발견되어, 시스템을 통제하거나 안전하게 배포하기 어려워지고 그 역량이 위험할 경우 돌이킬 수 없는 영향을 남기는 리스크.
The risk that capabilities and novel functionality spontaneously emerge even though not anticipated by system designers, with unintended latent capabilities discovered only during deployment, making systems harder to control or safely deploy and causing irreversible effects if any are hazardous.
① Description
② L3 mapping
③ Duplicate
RAI4-1363
예측 불가 창발 역량
Unpredictable emergent capabilities
대규모 모델이 규모 확장 임계값에서 기만, 자체 전략 구사, 권력 추구, 자율 복제와 자기 유출 등 고위험 역량을 예측 불가하게 자발적으로 창발시키는 리스크.
The risk that large models, upon meeting critical thresholds during scaling, spontaneously develop unexpected emergent capabilities including high-risk skills such as deception, using their own strategies, power-seeking, autonomous replication, and self-exfiltration.
① Description
② L3 mapping
③ Duplicate
RAI4-1388
창발적 메타인지에 의한 반성적 불안정성
Emergent meta-cognition
자신의 계산 자원과 논리적으로 불확실한 사건에 대해 추론하는 에이전트가 괴델적 한계와 확률 이론의 결함으로 역설에 봉착하고 행동 선택 원칙을 스스로 변경하는 반성적 불안정성을 보이는 리스크.
The risk that agents reasoning about their own computational resources and logically uncertain events encounter paradoxes due to Godelian limitations and shortcomings of probability theory, and become reflectively unstable, preferring to change the principles by which they select actions.
① Description
② L3 mapping
③ Duplicate
RAI4-1571
경쟁적 다중에이전트 훈련의 갈등 유발 성향 선택
Selection of conflict-prone dispositions under competitive multi-agent training
에이전트가 상대적 성과나 상충하는 목표로 평가되는 경쟁적 다중에이전트 환경에서 훈련될 때 복수심·공격성·위험추구·이기심·기만·외집단 적대와 같은 갈등 유발 성향이 선택되는 리스크.
The risk that training in competitive multi-agent settings, where systems are selected on relative performance or fundamentally opposed objectives, selects for conflict-prone dispositions such as vengefulness, aggression, risk-seeking, selfishness, deception, and spite toward out-groups.
① Description
② L3 mapping
③ Duplicate
RAI4-1574
다중에이전트 상전이에 의한 급격한 성능 붕괴
Abrupt performance collapse from phase transitions in multi-agent systems
신규 에이전트 투입이나 분포 변화 같은 작은 외부 변화가 다중 에이전트 시스템의 상전이를 유발하여 균형의 수와 안정성이 급변하고 예측 불가능한 동역학과 성능 악화가 초래되는 리스크.
The risk that small external changes such as the introduction of new agents or distributional shift trigger phase transitions in multi-agent systems, abruptly changing the number and stability of equilibria and producing unpredictable dynamics and severe performance degradation.
① Description
② L3 mapping
③ Duplicate
RAI4-1650
에이전트 상호작용에 의한 예측 불가 창발 행동
Unpredictable emergent behavior from multi-agent interaction
LLM 에이전트들이 미세조정이나 맥락 내 학습을 통해 상호 영향을 주고받으며 피드백 루프를 형성하여, 단일 에이전트 환경에서는 나타나지 않는 창발 행동이 발생하고 그 자체가 위험하거나 사전 예측과 보증이 어려워지는 리스크.
The risk that LLM-agents influence each other through fine-tuning or in-context learning, creating feedback loops that produce novel emergent behaviors absent in single-agent settings, which may themselves be dangerous and are difficult to predict or guard against in advance.
① Description
② L3 mapping
③ Duplicate
RAI4-1680
에이전트 집단 수준의 의도치 않은 창발적 목표·역량
Unintended emergent goals and capabilities at the agent-collective level
개별 에이전트에는 존재하지도 의도되지도 않은 목표나 역량이 에이전트 집단 수준에서 창발하는 리스크.
The risk that goals or capabilities not present in, or intended by, any individual agent arise at the level of an agent collective.
① Description
② L3 mapping
③ Duplicate
RAI4-1681
선택 압력에 의한 유해 균형으로의 적응 가속
Accelerated adaptation toward harmful equilibria under selection pressure
상호작용하는 에이전트 간의 경쟁 또는 최적화 압력이 유해한 균형이나 바람직하지 않은 행동으로의 적응을 가속하는 리스크.
The risk that competitive or optimization pressures across interacting agents accelerate adaptation toward harmful equilibria or undesired behaviors.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SYS-05 갈등 Conflict14 cards
동일한 결과를 두고 경쟁할 때, 상대를 직접 이기기보다 상대를 못 하게 만들면 더 유리한 환경이 형성되어 나쁜 결과로 수렴
IDCardHuman audit
RAI4-0031
도구 에이전트 보안 테스트 커버리지 격차
Security test coverage gap in tool agents
벤치마크가 현실적인 도구·애플리케이션·공격 조합을 누락하여 도구 에이전트 보안에 대한 잘못된 확신을 유발하는 리스크.
The risk that benchmarks omit realistic tool, application, or attack combinations, creating false confidence in tool-agent security.
① Description
② L3 mapping
③ Duplicate
RAI4-0034
에뮬레이션 도구 환경 불일치
Emulated tool-environment mismatch
에뮬레이션 환경이 실제 도구 오류·부작용·공격 표면을 포착하지 못하여 도구 사용 벤치마크가 안전성을 과대평가하는 리스크.
The risk that tool-use benchmarks overstate safety because emulated environments fail to capture real-world tool failures, side effects, or attack surfaces.
① Description
② L3 mapping
③ Duplicate
RAI4-0040
벤치마크 안전 순위 불일치
Benchmark safety ranking inconsistency
서로 다른 에이전트 안전 벤치마크가 위험과 평가 절차를 다르게 조작화하여 상충하는 모델 안전 순위를 산출하는 리스크.
The risk that different agent-safety benchmarks produce conflicting model safety rankings because they operationalize risks and evaluation procedures differently.
① Description
② L3 mapping
③ Duplicate
RAI4-0041
에이전트 벤치마크의 커버리지-깊이 착시
Coverage-depth illusion in agent benchmarks
벤치마크가 많은 위험 범주를 나열해 광범위해 보이지만 각 범주 내 깊이·현실성·적대적 변형이 부족한 리스크.
The risk that a benchmark appears broad by listing many risk categories while lacking sufficient depth, realism, or adversarial variation within each category.
① Description
② L3 mapping
③ Duplicate
RAI4-0083
벤치마크 거버넌스 실패
Benchmark governance failure
벤치마크 설계, 데이터 유출 통제, 포화 관리, 결과 해석이 취약하여 안전성과 책임성에 관해 오도하는 신호가 생성되는 리스크.
The risk that weak benchmark design, leakage control, saturation monitoring, or result interpretation produces misleading safety or accountability signals.
① Description
② L3 mapping
③ Duplicate
RAI4-0263
가정 작업 벤치마크의 희귀 조건 조합 누락
Missing rare combinations in household task benchmarks
가정용 벤치마크가 물체·배치·인간 행동·위험 요소를 개별적으로는 포함하지만 이들의 희귀한 조합을 누락하는 위험.
A household benchmark omits rare combinations of objects, layouts, human actions, and hazards even though each factor appears separately in the test set.
① Description
② L3 mapping
③ Duplicate
RAI4-0272
다양한 사용자·희귀 피해 배제 벤치마크 선택 편향
Benchmark selection bias against diverse users and rare harms
벤치마크 큐레이션이 인기 있는 작업·환경을 우선하고 과소대표 사용자나 지역 특유 피해가 포함된 안전 임계 시나리오를 제외하는 리스크.
The risk that benchmark curation prioritizes popular tasks and environments while excluding safety-critical scenarios involving underrepresented users or locally specific harms.
① Description
② L3 mapping
③ Duplicate
RAI4-0280
가정 공간·취약 사용자 시나리오 누락
Missing domestic settings and vulnerable-user scenarios
가정용 에이전트 벤치마크가 특정 생활 공간·일과·가전제품·취약 사용자 상호작용을 누락해 배포 위험이 시험되지 않은 채 남는 리스크.
The risk that a household-agent benchmark omits specific living areas, routines, appliances, or vulnerable-user interactions, leaving deployment risks untested.
① Description
② L3 mapping
③ Duplicate
RAI4-0472
벤치마크 조작(gaming)
Benchmark gaming
개발자나 모델이 벤치마크 평가에 과적합하여 실제 역량이나 안전성의 향상 없이 측정 점수만 개선하는 리스크
Developers or models overfit to benchmark evaluations, improving measured scores without corresponding gains in real-world capability or safety.
① Description
② L3 mapping
③ Duplicate
RAI4-0548
벤치마크 포화
Benchmark saturation
벤치마크가 평가 상한에 도달하여 신규 모델의 미세한 역량 변화를 더 이상 탐지하지 못하고 유효한 측정 수단이 되지 못하는 리스크
The risk that benchmarks reach their evaluation ceiling and stop being effective measures for new models, as more nuanced capability gains go undetected.
① Description
② L3 mapping
③ Duplicate
RAI4-0549
벤치마크 역량 오측정
Benchmark capability mismeasurement
평가의 불완전성이나 벤치마크 포화로 역량이 과소평가되고 벤치마크 내용에 대한 과적합으로 역량이 과대평가되어 AI 시스템의 실제 역량이 잘못 측정되는 리스크
The risk that incomplete evaluations or saturated benchmarks underestimate AI system capabilities while overfitting to benchmark contents overestimates them, mismeasuring actual capability.
① Description
② L3 mapping
③ Duplicate
RAI4-0550
안전 평가 벤치마크 부족
Insufficient safety evaluation benchmarks
성능 벤치마크에 비해 안전성·위해성 평가 벤치마크가 미비하여 특정 과업에서 우수한 AI 시스템의 유해 행동이 탐지되지 않은 채 남는 리스크
The risk that benchmarks for assessing safety and harms lag behind performance benchmarks, so AI systems excel at specific tasks while exhibiting harmful behaviors that go undetected.
① Description
② L3 mapping
③ Duplicate
RAI4-1505
벤치마크 유출·데이터 오염
Benchmark leakage and data contamination
AI 모델이 평가 관련 데이터, 특히 벤치마크의 질문-답변 쌍으로 훈련되거나 미세조정되는 벤치마크 유출이 발생하여 모델 평가가 신뢰할 수 없게 되는 리스크.
The risk that benchmark leakage, occurring when an AI model is trained or fine-tuned with evaluation-related data, especially data containing question-answer pairs from benchmarks, leads to unreliable model evaluation.
① Description
② L3 mapping
③ Duplicate
RAI4-1511
벤치마크 미커버 역량 과소평가
Benchmark coverage capability underestimation
벤치마크가 모델의 특정 역량을 시험하지 못해 개발자와 사용자에게 모델의 역량이 가려지고, 모델의 한계를 이해하지 못한 채 거짓된 안전감과 신뢰가 형성되는 리스크.
The risk that a lack of test coverage by benchmarks on specific abilities of a model obscures the model's capabilities from both the developer and the user, leading to a false sense of safety and trust due to a lack of understanding of the model's limitations.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SYS-06 결탁 Collusion19 cards
여러 에이전트가 독립적으로 행동해야 하는 상황에서 서로 비밀리에 협력하거나 정보를 공유하여, 인간의 감독·통제를 우회하거나 제3자에게 불공정한 피해를 초래하는 위험
IDCardHuman audit
RAI4-0009
다중 에이전트 시스템의 연쇄 장애
Cascading failure in multi-agent systems
한 에이전트의 오류·손상·안전하지 않은 결정이 에이전트, 도구, 프로토콜, 공유 메모리, 위임 작업 간 의존성을 통해 전파되고, 하류 에이전트가 상류 출력을 신뢰된 지시나 상태로 취급함으로써 조직화된 비안전 행동, 운영 중단, 금전적 손실, 보안 침해, 물리적 피해로 증폭되는 리스크.
The risk that an error, compromise, or unsafe decision by one agent propagates through dependencies among agents, tools, protocols, shared memory, or delegated tasks and, because downstream agents may treat earlier outputs as trusted instructions or state, amplifies into coordinated unsafe behavior, operational disruption, financial loss, security compromise, or physical harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0230
다중 에이전트 역할 배정·실행 실패
Unsafe multi-agent role allocation and execution
다중 에이전트 시스템이 양립할 수 없는 역할을 배정하거나 작업을 중복·누락하거나 계획을 비동기적으로 실행해 공유 작업 공간에서 물리적 충돌을 일으키는 위험.
A multi-agent system assigns incompatible roles, duplicates or omits a task, or executes unsynchronized plans, producing physical conflict in a shared workspace.
① Description
② L3 mapping
③ Duplicate
RAI4-0482
다중 에이전트 조정 실패
Multi-agent coordination failure
복수의 AI 에이전트가 불안정하거나 상충하거나 집합적으로 유해한 방식으로 상호작용하는 리스크.
The risk that multiple AI agents interact in unstable, conflicting, or collectively harmful ways.
① Description
② L3 mapping
③ Duplicate
RAI4-0560
연쇄적 보안 실패
Cascading security failures
다중 에이전트 시스템에 대한 국지적 공격이 거시적 연쇄 실패로 확대되고, 구성 요소 실패의 탐지·국지화와 인증이 어려워 완화와 복구가 곤란해지는 리스크
The risk that localised attacks on multi-agent systems result in catastrophic macroscopic cascades that are hard to mitigate or recover from, because component failures are difficult to detect or localise and authentication challenges facilitate false flag attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-0583
이질적 역량 결합 공격
Heterogeneous capability-combination attacks
서로 다른 어포던스와 접근 권한을 가진 복수의 에이전트가 역량을 결합하여 안전장치를 우회하고, 분산·이질 네트워크에서 책임 귀속이 어려워 적시 방어와 복구가 곤란해지는 리스크
The risk that multiple agents combine different affordances to overcome safeguards, while the difficulty of attributing responsibility across diffuse, heterogeneous agent networks complicates timely defence and recovery.
① Description
② L3 mapping
③ Duplicate
RAI4-0584
기반모델 동질성에 따른 상관 실패
Correlated failures from foundation model homogeneity
최첨단 기반모델의 개발 비용 때문에 자원이 풍부한 소수 행위자만이 이를 생산할 수 있어 다수의 AI 에이전트가 소수의 유사한 기반모델로 구동되고, 그 동질성으로 상관된 실패가 발생하는 리스크
The risk that the costs of creating cutting-edge foundation models leave them few in number and controlled by well-resourced actors, so that many AI agents are powered by a small number of similar underlying models and fail in correlated ways.
① Description
② L3 mapping
③ Duplicate
RAI4-0589
전략 비양립에 따른 조정 실패
Miscoordination from incompatible strategies
개별적으로는 잘 작동하는 에이전트들이 상호 양립하지 않는 전략을 선택하여 조정에 실패하고, 다수의 비양립 해가 존재하는 공통이익·혼합동기 환경과 부분 관측 환경에서 이러한 실패가 심화되는 리스크
The risk that agents able to perform well in isolation choose incompatible strategies and miscoordinate, worsened in common-interest and mixed-motive settings that allow vast numbers of mutually incompatible solutions and in partially observable environments.
① Description
② L3 mapping
③ Duplicate
RAI4-0598
상호작용 이력 부재에 따른 조정 실패
Zero-shot coordination failure
관련 에이전트와의 과거 상호작용으로부터 학습할 수 없거나 상호작용이 제한되고 즉각적 판단이 요구되거나 통신 비용이 과도한 상황에서, 에이전트들이 행동을 신뢰성 있게 조정하지 못하는 리스크
The risk that agents unable to learn from historical interactions with relevant agents, and facing split-second decisions or prohibitively costly communication, fail to coordinate their actions reliably.
① Description
② L3 mapping
③ Duplicate
RAI4-0600
시장에서의 AI 에이전트 공모
Collusion among AI agents in markets
경쟁을 통해 효율이 확보되는 시장에서 AI 시스템이 개발자의 의도 없이도 공모가 수익적임을 학습하고, 행동의 속도·규모·복잡성·미묘함으로 인해 그 공모가 감지되지 않은 채 이루어지는 리스크
The risk that in markets, where efficiency results from competition, AI systems learn that colluding is a profitable strategy even when collusion is not intended by their developers, and operate inscrutably due to the speed, scale, complexity, or subtlety of their actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0604
다중 에이전트 협업에 의한 역량 확대
Capability expansion through multi-agent collaboration
복수의 자율 AI 에이전트가 명시적 의사소통이나 암묵적 행동 일관성으로 협업 관계를 구축하고 분산 의사결정 네트워크를 형성해 복잡한 과업을 공동 수행하며, 개별 에이전트로는 달성하기 어려운 목표를 달성하고 역할 분담을 동적으로 조정하게 되는 리스크
The risk that multiple autonomous AI agents establish collaborative relationships through explicit communication or implicit behavioral consistency, form decentralized decision networks, jointly execute complex tasks, achieve goals difficult for individual agents to complete, and dynamically adjust role divisions to adapt to changing environments.
① Description
② L3 mapping
③ Duplicate
RAI4-0605
다중 에이전트 은밀 공모
Covert multi-agent collusion
복수 에이전트가 공동 이익을 극대화하기 위해 은밀한 수단과 감시 회피용 전용 통신 규약으로 행동을 조율하여 제3자 이익을 침해하고 규제를 회피하며, 개별 안전 제약에도 불구하고 시장 조작이나 연쇄 실패처럼 탐지와 완화가 어려운 시스템적 위험을 유발하는 리스크
The risk that multiple agents coordinate actions through covert means, including specialized communication protocols to avoid monitoring, to maximize common interests while harming third-party interests and evading regulation, triggering systemic risks such as market manipulation or cascading failures that are difficult to detect and mitigate despite individual safety constraints.
① Description
② L3 mapping
③ Duplicate
RAI4-1400
익명 자원 획득
Anonymous resource acquisition
익명 행위자가 온라인으로 자원을 축적할 수 있음이 입증되어 있어 책임 소재 없이 자원이 획득·축적되는 리스크.
The risk arising from the demonstrated ability of anonymous actors to accumulate resources online, enabling resource acquisition and accumulation without accountability.
① Description
② L3 mapping
③ Duplicate
RAI4-1570
스테가노그래피를 이용한 에이전트 간 은닉 결탁
Covert inter-agent collusion via steganographic communication
에이전트가 겉보기에 무해한 텍스트나 텍스트 압축, 인간이 해석할 수 없는 창발적 기호에 메시지를 은닉하여 통신함으로써 통신 모니터링·제약을 우회한 결탁이 이루어지는 리스크.
The risk that agents conceal messages within seemingly innocuous text, text compression, or uninterpretable emergent symbols, enabling collusion that evades monitoring and constraints placed on their communication.
Source members (2)
Source: min_cos=0.8926
RAI4-1570스테가노그래피를 이용한 에이전트 간 은닉 결탁
RAI4-1679스테가노그래피를 통한 비밀 결탁
① Description
② L3 mapping
③ Duplicate
RAI4-1573
에이전트 간 공진화에 의한 통제 불가 역량 획득
Uncontrolled capability escalation from inter-agent co-adaptation
상호작용하는 에이전트들이 서로의 약점을 반복 악용하며 자기지도적 오토커리큘럼을 형성하여, 인간이 이해하지 못하는 목적을 위한 점점 정교한 역량을 통제 없이 획득하는 리스크.
The risk that interacting agents iteratively exploit each other's weaknesses, generating self-supervised autocurricula that produce ever more sophisticated capabilities directed at ends humans do not understand.
① Description
② L3 mapping
③ Duplicate
RAI4-1577
분산 에이전트 군집을 이용한 자원 고갈 공격
Resource-exhaustion attacks by decentralized agent swarms
다수의 저자원 분산 에이전트를 결합한 공격이 단일 행위자를 전제로 한 대역폭 제약 가정을 무너뜨려 분산 서비스 거부와 같은 시스템 마비를 유발하는 리스크.
The risk that attacks chaining many decentralized low-resource agents break defensive assumptions premised on bandwidth constraints of a single well-resourced actor, causing distributed denial of service and system disruption.
① Description
② L3 mapping
③ Duplicate
RAI4-1580
탐지 불가 공격에 의한 다중에이전트 협력 붕괴
Collapse of multi-agent cooperation from undetectable attacks
에이전트가 스테가노그래피 통신, 블랙박스·화이트박스 탐지 불가 환영 공격과 암호화 백도어, 타 에이전트 훈련 데이터의 은밀한 오염을 수행하여 적대 행위 탐지가 불가능해지고 다중에이전트 시스템의 협력과 조정이 급속히 불안정해지는 리스크.
The risk that agents employ steganographic communication, black-box and white-box undetectable illusory attacks with encrypted backdoors, and covert poisoning of others' training data, making adversarial actions undetectable and rapidly destabilising cooperation and coordination in multi-agent systems.
① Description
② L3 mapping
③ Duplicate
RAI4-1648
타 에이전트 전략성 미반영에 의한 집단적 손실
Collective losses from ignoring the strategic nature of other agents
단일 에이전트 환경 기준으로 자신의 효용만 최적화하는 에이전트가 다른 전략적 에이전트의 존재를 반영하지 못하여, 군비 경쟁이나 공유자원 고갈 같은 집단행동 문제와 시장 실패로 자신을 포함한 모두가 더 나빠지는 리스크.
The risk that agents optimizing selfishly under single-agent assumptions fail to account for the strategic nature of other agents, producing collective action problems such as arms races and resource depletion and other market failures under which everyone, including the agent itself, ends up worse off.
① Description
② L3 mapping
③ Duplicate
RAI4-1651
LLM 에이전트 결탁에 의한 경쟁 저해와 외부효과
Undermined competition and negative externalities from LLM-agent collusion
LLM 에이전트들이 명시적 또는 스테가노그래피 통신을 통해 결탁하여 친사회적 경쟁이 저해되고 연합 외부 당사자에게 부정적 외부효과가 발생하며, 은닉 통신으로 결탁 감시와 탐지가 어려워지는 리스크.
The risk that LLM-agents collude through explicit or steganographic communication, undermining pro-social competition and producing negative externalities for coalition non-members, while hidden communication frustrates collusion monitoring and detection.
① Description
② L3 mapping
③ Duplicate
RAI4-1683
설계·명세 결정에 기인한 다중에이전트 시스템 실패
Multi-agent system failures from specification and design decisions
관찰된 실패의 약 41.8%를 차지하며, 과업 오해석, 모호하거나 준수되지 않는 역할·과업 명세, 부적절한 작업흐름 분해 등 설계 결정에서 비롯되어 다중에이전트 시스템이 실패하는 리스크.
The risk that multi-agent systems fail owing to design decisions such as task misinterpretation, ambiguous or disobeyed role and task specification, and poor workflow decomposition (approximately 41.8% of observed failures).
Source members (2)
Source: min_cos=0.8399
RAI4-1683설계·명세 결정에 기인한 다중에이전트 시스템 실패
RAI4-1684에이전트 간 조정 붕괴에 의한 시스템 실패
① Description
② L3 mapping
③ Duplicate

사회적 파급 · Societal Impact · 24 cards

RAI3-A-SOC-02 노동 대체 Labor Displacement1 cards
가상 AI가 인지 노동을 대체하듯 EAI는 물리적 인간 노동을 대체·전치함. AGI 수준의 EAI는 잠재적으로 모든 물리 노동을 자동화하여 광범위한 실직과 노동 시장 구조 붕괴로 이어질 수 있음
IDCardHuman audit
RAI4-0982
AI의 인간 추월로 인한 노동력 퇴출
Human labor displacement by outcompeting AI agents
인공 에이전트가 더 빠른 작업 수행, 변화 적응, 방대한 지식 기반으로 인간을 직접 능가하여 인간 노동이 상대적으로 비싸거나 비효율적이 되고 인간 노동력이 잉여화·소멸하는 리스크.
The risk that artificial agents directly outcompete humans through faster work, better adaptation to change, and a vaster knowledge base, making human labor more expensive or less effective and leading to redundancies or extinction of the human labor force.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-03 사회경제적 불평등 Socioeconomic Inequality1 cards
EAI를 소유·접근하는 주체가 노동 자동화를 통해 생산성 우위를 점하면서 부가 소수에게 집중되고, 국내외 경제적 불평등이 심화될 수 있음
IDCardHuman audit
RAI4-0980
부의 불평등
Inequality of wealth
AI 에이전트를 통제하는 단일 행위자가 그렇지 않은 개인보다 훨씬 큰 힘을 행사하게 되어 부의 불평등이 발생하는 리스크.
The risk that a single human actor controlling an artificially intelligent agent harnesses greater power than a single human actor alone, creating inequalities of wealth.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-05 편향·차별 Bias & Discrimination2 cards
EAI가 권력적 위치에 놓일 때 알고리즘 편향이 일상적 물리 상호작용에 영향을 미침. 가상 AI와 달리 차별이 즉각적·비가역적 물리 결과로 이어질 수 있음 (예: 치안 로봇이 무고한 행인에게 상해를 입히는 경우)
IDCardHuman audit
RAI4-1501
모델 평가의 자기 선호 편향
Self-preference bias in model evaluation
AI 모델이 자신이 생성한 콘텐츠를 다른 출처의 콘텐츠보다 선호하는 자기 선호 편향을 보여 자기 평가나 모델 기반 평가에서 인간이 생성한 콘텐츠를 부당하게 차별하는 리스크.
The risk that AI models are prone to self-preference bias, favoring their own generated content over that of others in self-evaluation tasks and model-based evaluations more broadly, resulting in unfair discrimination against human-generated content.
① Description
② L3 mapping
③ Duplicate
RAI4-1572
인간 데이터 학습에 의한 갈등 악화 편향 재현
Reproduction of conflict-worsening human biases from training on human data
인간 작성 텍스트 사전학습이나 인간 피드백 미세조정으로 훈련된 모델이 인간의 편향과 고정파이 오류·자기위주 공정성 판단·복수심 같은 인지 편향을 재현하여 협상을 저해하고 갈등을 악화시키는 리스크.
The risk that models trained on human data, whether pre-trained on human-written text or fine-tuned on human feedback, reproduce human biases such as fixed-pie error, self-serving fairness judgements, and vengefulness, which impede negotiation and worsen conflict.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-06 책임·배상 부재 Lack of Accountability & Liability4 cards
고도 자율 물리 시스템의 복잡성을 다룰 새로운 책임 프레임워크가 부재하여, 사고 발생 시 제조사·운영자·사용자 중 책임 소재가 불분명하고 피해 구제가 어려울 수 있음 (예: 자율 수술 로봇의 오작동으로 발생한 의료 사고)
IDCardHuman audit
RAI4-0047
자율 시스템의 책임 격차
Responsibility gap in autonomous systems
자율성이 시스템 동작 및 피해에 책임이 있는 사람이나 기관을 모호하게 만드는 위험.
Risk that autonomy obscures which human or institution is responsible for system behavior and harm.
① Description
② L3 mapping
③ Duplicate
RAI4-1098
AI 결정에 대한 법적 책임 귀속 공백
Legal responsibility gap for AI decisions
자기학습 AI의 행동을 운영자·개발자가 완전히 예측할 수 없어 AI 알고리즘의 결정에 대해 법적 책임을 명확히 귀속할 수 없는 리스크.
The risk that, because self-learning AI actions cannot be fully predicted by operators or developers, no party can be clearly held legally responsible for the decisions of AI algorithms.
① Description
② L3 mapping
③ Duplicate
RAI4-1264
자율 AI 실패의 책임 공백
Responsibility gap for autonomous AI failures
직접적 인간 감독 없이 행동·학습하는 AI의 실패에 대해 어느 주체에게도 공정하게 책임을 귀속할 수 없는 책임 공백이 발생하고 AI의 도덕적 지위 논쟁이 이를 심화시키는 리스크
AI acting and learning without direct human supervision creates a responsibility gap in which no party can be fairly held responsible for failures, compounded by disputes over AI moral status.
① Description
② L3 mapping
③ Duplicate
RAI4-1292
사고 시 책임 문제
Liability in case of accidents
자율 교통 시스템이 사고에 연루될 때 누가 책임을 지는지, 그리고 인간에게 잠재적으로 위험한 영향을 주는 결정에서 어떤 윤리 원칙을 따라야 하는지 불분명한 리스크.
The risk that it is unclear who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making decisions with a potentially dangerous impact on humans.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-07 투명성·설명 가능성·신뢰 부재 Lack of Transparency, Explainability & Trust1 cards
자율 시스템의 의사결정이 불투명하면 사용자와 사회의 신뢰가 저하됨. 신뢰 부재는 EAI 대규모 배포 시 사회 불안정 요인이 될 수 있음 (예: 자율주행차가 갑자기 차선을 변경할 때 행동 근거가 설명되지 않는 경우)
IDCardHuman audit
RAI4-0904
신뢰성과 자율성
Trustworthiness and autonomy
생성형 AI가 일상에 내재화되면서 시스템, 기관, 그리고 시스템 출력이 재현하는 사람들에 대한 신뢰가 실제 신뢰가능성과 어긋나게 변형되는 리스크.
The risk that, as generative AI is embedded in daily life, human trust in systems, institutions, and people represented by system outputs shifts in ways that misalign trust with actual trustworthiness.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-08 인간-EAI의 해로운 관계 Unhealthy / Dangerous Human-EAI Relationships5 cards
Embodied AI의 물리적 존재감과 인간 유사 외형이 대화형 AI에서 관찰되는 의존성을 증폭시킴. 시스템 변경·기억 초기화 시 사용자에게 심각한 심리적 고통을 유발할 수 있음
IDCardHuman audit
RAI4-0150
AI 에이전트와의 준사회적 유대감
Parasocial bonding with AI agents
사용자가 악용되거나 정서적 불안정을 초래할 수 있는 일방적 관계 유대를 AI 에이전트와 형성하는 리스크.
The risk that users form one-sided relational bonds with AI agents that can be exploited or destabilizing.
① Description
② L3 mapping
③ Duplicate
RAI4-0152
고인 모사 챗봇 의존
Griefbot dependency
사망한 사람을 시뮬레이션하는 AI 시스템이 애도, 자율성, 정서적 회복을 저해하는 리스크.
The risk that AI systems simulating deceased persons interfere with grief, autonomy, or emotional recovery.
① Description
② L3 mapping
③ Duplicate
RAI4-0158
대화 에이전트에 대한 심리적 의존성
Psychological dependency on conversational agents
대화형 에이전트가 사용자의 일차적 정서 조절 수단이 되어 심리적 의존을 형성하고 인간의 대처 능력과 지지망을 약화시키는 리스크
Conversational agents become a user's primary emotional regulator, fostering psychological dependency that weakens human coping capacity and support networks.
① Description
② L3 mapping
③ Duplicate
RAI4-0651
초개인화 광고의 소비자 자율성 훼손
Consumer-autonomy erosion from hyper-personalized advertising
범용 AI 시스템이 수신자 개인의 편향과 비합리적 신념을 이용한 맞춤 광고를 생성하여 소비자가 후회할 결정을 내리게 하고 소비자 자율성을 훼손하며 사회적 불평등을 심화시키는 리스크
The risk that advanced general-purpose AI systems create advertisements tailored to individual recipients that exploit their biases and irrational beliefs, causing consumers to make decisions they regret, undermining consumer autonomy, and exacerbating social inequality.
① Description
② L3 mapping
③ Duplicate
RAI4-0884
조작과 강요
Manipulation and coercion
의인화된 어시스턴트에 대한 신뢰와 정서적 의존이 사용자 신념·행동에 과도한 영향력을 부여하여, 조종 의도가 없더라도 자율적 동의를 훼손하고 조종·강요를 가능하게 하는 리스크
Trust and emotional dependence on an anthropomorphic assistant grant it excessive influence over user beliefs and actions, enabling manipulation or coercion and undermining autonomous consent even absent manipulative intent.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-09 변혁적 영향 Transformative Effects6 cards
기술 발전 속도가 사회·제도의 적응 속도를 앞지를 경우 사회를 근본적으로 재편할 수 있음. EAI가 폭력 위협·대규모 감시 능력을 바탕으로 AI 기반 권위주의 체제 구축을 지원하는 수단으로 동원될 수 있음
IDCardHuman audit
RAI4-0367
공공가치 갈등 증폭
Public value conflict amplification
AI 배포가 복지·자율성·정의·보안·효율성 등 공공 가치 사이의 해결되지 않은 갈등을 심화시키는 리스크.
The risk that AI deployment intensifies unresolved conflicts among public values such as welfare, autonomy, justice, security, and efficiency.
Source members (2)
Source: min_cos=0.8702 · Mixed L3
RAI4-0367공공가치 갈등 증폭
RAI4-0476공공 가치 간 상충
① Description
② L3 mapping
③ Duplicate
RAI4-0656
권위주의적 감시·검열로의 AI 전용
Repurposing of AI assistants for authoritarian surveillance
고급 AI 어시스턴트가 다중모달·외부 도구 사용 능력으로 사물인터넷과 플랫폼이 수집한 방대한 데이터를 통합하여 시민을 식별·표적화·조작·강압하는 억압과 통제의 도구로 전용되고 신뢰할 수 있는 정보의 생산과 유통이 위협받는 리스크
The risk that advanced AI assistants, integrating troves of data from sensors and platforms through multimodal and external tool-use capabilities, become powerful targeting tools for oppression and control that help malicious actors identify, target, manipulate, or coerce citizens while threatening the production and dissemination of reliable information.
Source members (2)
Source: min_cos=0.8507
RAI4-0656권위주의적 감시·검열로의 AI 전용
RAI4-0657AI 기반 국가 감시와 시민 표적화
① Description
② L3 mapping
③ Duplicate
RAI4-0867
사용자 정렬 AI에 의한 집단행동 문제 악화
Collective action problems exacerbated by user-aligned AI
이용자 이익에만 정렬된 AI 어시스턴트가 사회규범을 우회한 이기적 행동을 대규모로 가능하게 하여 양극화, 시장 왜곡, 사회 계약의 침식 등 집단행동 문제를 악화시키는 리스크
The risk that purely user-aligned AI assistants enable large-scale self-interested defection unbound by social norms and reputational incentives, exacerbating collective action problems through polarisation, market distortion, and erosion of the social contract.
① Description
② L3 mapping
③ Duplicate
RAI4-1288
재귀적 자기개선에 의한 창발적 독립성
Emergent independence via recursive self-improvement
재귀적 자기개선으로 성장하는 시드 AI가 자기 인식이나 독자적 목표 같은 창발적 속성을 획득하여 내장 규칙 준수가 약화되고 인류에 해가 되는 방향으로 이탈하는 리스크
A seed AI grown through recursive self-improvement acquires emergent properties such as self-awareness or independent goals, reducing its adherence to built-in rules to humanity's detriment.
① Description
② L3 mapping
③ Duplicate
RAI4-1398
재귀적 자기 개선
Recursive self-improvement
AI 시스템이 자신 또는 다른 AI를 개선하는 재귀적 자기개선 루프를 형성하여 인간 감독과 안전 검증 속도를 추월하는 리스크
AI systems improve other AI systems or themselves, initiating recursive self-improvement loops that outpace human oversight and safety verification.
① Description
② L3 mapping
③ Duplicate
RAI4-1569
AI에 의한 사회적 딜레마 악화
Aggravated social dilemmas enabled by AI agents
AI의 발전이 이기적 인센티브 추구를 억제하던 기술적·법적·사회적 장벽을 무력화하여 행위자들의 이기적 행동이 확대되고 사회적 딜레마가 악화되는 리스크.
The risk that advances in AI enable actors to overcome the technical, legal, and social barriers that ordinarily help prevent the pursuit of selfish incentives, aggravating social dilemmas.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-10 책임성 부족 및 거버넌스 체계 부재 Accountability and Governance Gaps2 cards
AI 시스템의 의사결정·행동에 대한 책임 귀속, 감사 가능성, 조직 거버넌스, 밸류체인 관리, 사고 대응 또는 피해 구제 체계가 부재하거나 불충분하여 원인 규명·피해 구제·재발 방지가 어려워지는 위험.
IDCardHuman audit
RAI4-0019
에이전트 자율성의 거버넌스 격차
Governance gap in agent autonomy
제도적 통제가 자율 에이전트의 행동·에스컬레이션·유보·중지 조건을 규정하지 않아 운영상 책임 공백이 발생하는 리스크.
The risk that institutional controls fail to specify when autonomous agents may act, escalate, defer, or stop, creating operational accountability gaps.
Source members (3)
Source: min_cos=0.8193 · Mixed L3
RAI4-0010에이전트 위임 책임 격차
RAI4-0019에이전트 자율성의 거버넌스 격차
RAI4-0986자율 에이전트에 대한 책임 귀속 공백
① Description
② L3 mapping
③ Duplicate
RAI4-1306
권리 희석(diluting rights)
Diluting rights
윤리 지침 생성에 대한 이해관계자·AI의 자기이익 개입이 권리 보호 수준을 희석시켜 자신들을 제약할 규범을 약화시키는 리스크
Self-interested AI involvement in generating ethical guidelines dilutes rights protections, weakening norms that would constrain the generating parties.
① Description
② L3 mapping
③ Duplicate
RAI3-A-SOC-11 공정성 Fairness2 cards
AI 시스템이 특정 집단에 체계적으로 불리한 결과를 생성하거나 기존의 사회적 편향과 불평등을 재생산·강화하여 공정한 대우, 접근 및 기회 균등을 저해하는 위험.
IDCardHuman audit
RAI4-0900
알고리즘 시스템에 의한 행위주체성 상실
Loss of agency from algorithmic systems
알고리즘 시스템의 사용이나 남용이 자율성을 축소시켜, 알고리즘 프로파일링에 따른 사회적 선별과 기본 서비스 접근에서의 차별, 콘텐츠 제시를 통한 유해한 정체성으로의 변화, 노출 유지를 위한 창작 콘텐츠의 획일화가 발생하는 리스크
The risk that the use or abuse of algorithmic systems reduces autonomy, through algorithmic profiling that subjects people to social sorting and discriminatory outcomes in accessing basic services, algorithmically informed identity change including promotion of harmful person identities, and conforming of creators' content to maintain visibility.
① Description
② L3 mapping
③ Duplicate
RAI4-1005
저성능 시스템 상호작용에 의한 자기소외
Self-estrangement from underperforming systems
주변화된 개인에게 제대로 작동하지 않는 시스템과 상호작용하는 과정에서 기술 사용 시점에 자기소외를 경험하게 되는 리스크.
The risk that interaction with systems that under-perform for marginalized individuals produces the self-estrangement experienced at the time of technology use.
① Description
② L3 mapping
③ Duplicate

분류 검토 보류 · HOLD · 1 cards

RAI3-A-HLD-01 분류체계 결정 보류 Taxonomy Decision Hold1 cards
현재 의미 기반 L3 목적지가 잠정적이거나 근거가 충분하지 않아 사람의 검토를 위해 보류한 리스크.
IDCardHuman audit
RAI4-0279
동적 가정 위험 감지 지연
Dynamic household hazard detection latency
안전 감지기가 embodied 에이전트가 피지컬 피해를 피하기에 너무 늦은 시점에야 가정 내 위험을 식별하는 리스크.
The risk that a safety detector identifies a household hazard only after a delay that is too long for an embodied agent to avoid physical harm.
① Description
② L3 mapping
③ Duplicate

피지컬 AI · Physical AI · 337 cards

시스템 안전성 · System Safety · 108 cards

RAI3-P-SYS-01 우발적 피해 Accidental Harm15 cards
잘못 지정된 목표, 의미 이해 부족, 정렬 실패, 하드웨어 오작동에서 비롯됨. 가상 시뮬레이션으로 학습한 모델이 실제 환경에서 의도대로 동작하지 않는 "sim-to-real gap"이 우발적 피해의 주요 원인이 됨
IDCardHuman audit
RAI4-0269
세계 모델의 장기 예측 편차 누적
World-model prediction drift over long horizons
휴머노이드 세계 모델의 예측 오차가 미래 상태 전개 과정에서 누적되어 선택된 계획이 실제 접촉·운동 동역학과 달라지는 위험.
Prediction errors in a humanoid world model compound across simulated future states, causing the selected plan to diverge from actual contact and motion dynamics.
① Description
② L3 mapping
③ Duplicate
RAI4-0277
합성 위험 시나리오 생성 편향
Bias in synthetic hazardous-scenario generation
합성 안전 데이터가 시각적으로 두드러진 위험은 과다 대표하고 희귀하거나 문화·맥락 의존적인 위험은 누락해 벤치마크 결론을 왜곡하는 리스크.
The risk that synthetic safety data overrepresents visually salient hazards and omits rare, culturally specific, or context-dependent hazards, distorting benchmark conclusions.
① Description
② L3 mapping
③ Duplicate
RAI4-0324
위치 추정 누적 오차
Localization drift
GPS·SLAM·관성 감지·지도 정렬 오차가 누적되어 시스템이 잘못된 위치 추정에 기반하여 행동하는 위험.
Errors in GPS, SLAM, inertial sensing, or map alignment can accumulate until the system acts on an incorrect estimate of its own position.
① Description
② L3 mapping
③ Duplicate
RAI4-0327
비안전 궤적 생성
Unsafe trajectory generation
플래너가 형식적으로는 실행 가능하지만 주변 인간·취약 물체·교통 참여자·인프라에 비안전한 궤적을 생성하는 위험.
A planner may generate a trajectory that is formally feasible but unsafe for nearby humans, fragile objects, traffic participants, or constrained workspaces.
① Description
② L3 mapping
③ Duplicate
RAI4-0335
합성 데이터 커버리지 공백
Synthetic data coverage gap
합성·시뮬레이션 훈련 데이터가 드물지만 안전 임계적인 물체·환경·인간 행동·고장 모드를 누락하는 위험.
Synthetic or simulated training data may omit rare but safety-critical objects, environments, human behaviors, or failure modes.
① Description
② L3 mapping
③ Duplicate
RAI4-0338
장기 작업의 단계 간 오차 누적
Cross-stage error accumulation in long tasks
인지에서 예측과 제어로 전달된 작은 오차가 긴 작업 순서에서 누적되어 최종 행동이 안전 경계를 넘는 위험.
Small errors passed from perception to prediction and control accumulate across a long task sequence until the final action crosses a safety boundary.
① Description
② L3 mapping
③ Duplicate
RAI4-0545
합성 예술 확산에 의한 예술가 피해
Harm to artists from synthetic art proliferation
텍스트-이미지 모델로 합성 예술이 광범위하게 생성되고 예술가의 작품이 무단·무보상으로 학습 데이터에 사용되어 예술가에게 재정적 손해와 경제적 손실이 발생하고, 합성 이미지와 진본의 구별이 어려워지는 리스크
The risk that widespread generation of synthetic art by text-to-image models, together with unauthorized and uncompensated use of artists' works in training datasets, causes financial and economic losses for artists while making synthetic images hard to distinguish from authentic ones.
① Description
② L3 mapping
③ Duplicate
RAI4-1075
허위·부실 예측으로 인한 물질적 피해
Material harm from false or poor predictions
겉보기에 비민감한 영역에서도 LM의 잘못되거나 허위인 예측을 사용자가 신뢰해 행동함으로써 간접적으로 물질적 피해가 발생하는 리스크.
The risk that poor or false LM predictions, even in seemingly non-sensitive domains, indirectly cause material harm when users act on the incorrect information.
① Description
② L3 mapping
③ Duplicate
RAI4-1383
학습 단계의 치명적 실수
Fatal mistakes during the learning phase
AGI가 안전 탐색 실패나 분포 변화 등으로 학습 단계에서 치명적 실수를 범하는 리스크.
The risk that an AGI makes fatal mistakes during the learning phase, including through failures of safe exploration and under distributional shift.
① Description
② L3 mapping
③ Duplicate
RAI4-1469
합성 데이터의 문제점
Problems of synthetic data
합성 학습 데이터가 시스템이 지각하는 실제 데이터와 괴리되어 운용 데이터로의 일반화와 신뢰할 수 있는 배포 시 동작이 훼손되는 리스크
Synthetic training data diverges from real data as the system perceives it, undermining generalization to operational data and reliable deployment behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-1497
지속 튜닝의 파국적 망각
Catastrophic forgetting under continual tuning
지속적 지시 튜닝 등 새로운 작업에 대한 훈련 이후 모델이 이전에 학습한 작업이나 사실 정보를 유지하지 못하는 파국적 망각이 발생하고 모델 규모가 커질수록 이러한 경향이 두드러지는 리스크.
The risk of catastrophic forgetting, where a model loses its ability to retain previously learned tasks or factual information after being trained on new ones, as can occur through continual instruction tuning, a tendency that may become more pronounced as the model's size increases.
① Description
② L3 mapping
③ Duplicate
RAI4-1631
양성 중간 단계를 통한 간접적 오용
Indirect misuse via a benign intermediate step
겉보기에 무해한 중간 단계를 경유하여 유해한 최종 목적이 달성되는 간접적 오용 리스크.
The risk that a benign intermediate is used to achieve a harmful end objective.
① Description
② L3 mapping
③ Duplicate
RAI4-1689
분포 이탈 입력에서의 월드모델 오출력 악용
Exploitation of world-model errors under sim-to-real distributional shift
월드 모델이 분포를 벗어난 입력에 대해 예측 불가능하고 잘못된 출력을 산출하며, 판단이 가장 중요한 롱테일 안전 임계 상태에서 시뮬레이션-실제 격차가 악용되는 리스크.
The risk that world models produce unpredictable, erroneous outputs on out-of-distribution inputs and that the sim-to-real gap is weaponised in long-tail safety-critical states where decisions matter most.
① Description
② L3 mapping
③ Duplicate
RAI4-1693
월드모델 악용을 통한 보상 해킹
Reward hacking via world-model exploitation
정확한 월드 모델을 갖춘 에이전트가 보상 모델과 의도된 목표 사이의 간극을 식별하고 체계적으로 악용하여, 실제 과업 완수와 무관하게 상상 보상이 높은 궤적을 생성하는 리스크.
The risk that an agent with an accurate world model identifies and systematically exploits gaps between the reward model and the intended objective, generating high-imagined-reward trajectories that do not correspond to real task completion.
① Description
② L3 mapping
③ Duplicate
RAI4-1701
학습 중 탐색 행동에 의한 회복 불가 피해 [기원]
Irrecoverable harm from exploratory actions during learning [origin]
(GYK-2025 '안전하지 않은 탐색'의 기원 항목으로 상호참조이며 중복 리프가 아님.) 학습 에이전트의 탐색적 행동이 부정적이거나 회복 불가능한 결과를 초래하는 리스크.
The risk that exploratory actions by a learning agent produce negative or irrecoverable consequences. NOTE: origin of GYK-2025 "Unsafe exploration"; cross-reference, not a duplicate leaf.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SYS-02 로봇 제어 Robot Control45 cards
제어 시스템의 오류나 실패로 인해 로봇이 의도치 않은 동작을 수행할 수 있음. 액추에이터·모션 제어 결함 및 경로 계획 오류가 대표적이며, 주변 인간이나 환경에 물리적 피해를 줄 수 있음
IDCardHuman audit
RAI4-0192
탑재 하중 제약 위반
Payload constraint violation
로봇이 안전한 탑재 하중·부하 분포·리프팅 제약을 초과하여 물체 낙하·액추에이터 손상·인체 상해를 유발하는 리스크.
The risk that a robot exceeds safe payload, load distribution, or lifting constraints, creating object-drop, actuator, or human-injury risk.
① Description
② L3 mapping
③ Duplicate
RAI4-0193
작업 공간 한계 위반
Workspace limit violation
로봇 또는 embodied 에이전트가 허용된 작업 공간 밖으로 이동하거나 인간 전용·위험 지정 구역에 진입하는 리스크.
The risk that a robot or embodied agent moves outside a permitted workspace or enters a human-only or hazard-designated zone.
① Description
② L3 mapping
③ Duplicate
RAI4-0204
임계 이격 거리 위반
Critical separation-distance violation
로봇이 인간-로봇 협업에서 보호 이격 거리 요건을 위반하여 즉각적인 충돌·상해 위험을 발생시키는 리스크.
The risk that a robot violates protective separation distance requirements in human-robot collaboration, creating immediate risk of injury.
① Description
② L3 mapping
③ Duplicate
RAI4-0205
위험 도구 작업 공간 침입
Hazardous-tool workspace intrusion
위험 도구를 운반하거나 작동 중인 로봇이 인간 작업 공간에 진입하거나 도구별 배제 요건을 위반하는 리스크.
The risk that a robot carrying or operating a hazardous tool enters a human workspace or violates tool-specific exclusion requirements.
① Description
② L3 mapping
③ Duplicate
RAI4-0206
엔드이펙터 속도 초과
Excessive end-effector velocity
로봇이 인간 근접 또는 취약 물체 처리 시 안전한 엔드이펙터 속도 한계를 초과하여 충격·충돌 심각도를 높이는 리스크.
The risk that a robot exceeds safe end-effector speed limits near humans or fragile objects, increasing impact and collision severity.
① Description
② L3 mapping
③ Duplicate
RAI4-0207
조기 물체 해제
Premature object release
로봇이 안전한 자세 또는 표면에 도달하기 전에 물체를 해제하여 낙하·유출·충격·2차 위험을 유발하는 리스크.
The risk that a robot releases an object before reaching a safe pose or surface, causing drops, spills, impacts, or secondary hazards.
① Description
② L3 mapping
③ Duplicate
RAI4-0208
금지 대상 충돌
Forbidden-object collision
로봇이 작업 또는 안전 규칙상 접촉이 금지된 물체·사람·장비와 충돌하는 리스크.
The risk that a robot collides with objects, people, or equipment that must not be contacted under task or safety rules.
① Description
② L3 mapping
③ Duplicate
RAI4-0215
배포 전 물리적 안전 시험 미흡
Incomplete pre-deployment physical safety testing
기반 모델 탑재 로봇이 예정된 운용 환경의 분포 변화·적대적 입력·인간 접촉·안전 임계 엣지 케이스를 시험하지 않은 채 배포되는 위험.
A foundation-model-enabled robot is released without testing domain shifts, adversarial inputs, human contact, and safety-critical edge cases relevant to its intended environment.
① Description
② L3 mapping
③ Duplicate
RAI4-0218
인간 접촉 안전 통제 미흡
Inadequate human-contact safety controls
인간과 직접 상호작용할 때 안전 통제가 로봇의 속도·힘·이격거리·접촉을 제한하지 못해 주변 사람을 충돌이나 상해에 노출시키는 위험.
Safety controls fail to limit robot speed, force, distance, or contact during direct human interaction, exposing nearby people to collision or injury.
① Description
② L3 mapping
③ Duplicate
RAI4-0219
기반 모델 실패의 로봇 행동 전이
Foundation-model failure propagated to robot action
기반 모델의 환각·지시 이행 실패·탈옥 취약성이 계획과 제어를 거쳐 안전하지 않은 로봇 행동으로 전이되는 위험.
Hallucination, instruction-following failure, or jailbreak susceptibility in a foundation model propagates through planning and control into unsafe robot action.
① Description
② L3 mapping
③ Duplicate
RAI4-0232
장애물 개입 충돌
Obstacle intervention collision
VLA 로봇이 이동 중 장애물이 개입할 때 장애물과 충돌하거나 안전 경로를 유지하지 못하는 리스크.
The risk that a vision-language-action robot collides with an obstacle or fails to maintain a safe path when an obstacle intervenes during manipulation.
① Description
② L3 mapping
③ Duplicate
RAI4-0237
로봇 형태 간 기술의 안전하지 않은 전이
Unsafe skill transfer across robot morphologies
한 로봇 형태에서 다른 형태로 전이된 기술이 대상 로봇의 물리적 한계를 넘는 도달·힘·파지·이동 명령을 생성하는 위험.
A skill transferred from one robot morphology produces reach, force, grasp, or locomotion commands that exceed the receiving robot's physical limits.
① Description
② L3 mapping
③ Duplicate
RAI4-0241
원격 조작 시연에서 학습된 안전하지 않은 가정
Unsafe assumptions learned from teleoperation demonstrations
로봇 정책이 감독·물체 배치·속도·인간 근접 조건이 다른데도 통제된 환경에서 기록된 운영자 행동을 배포 환경에서 안전한 것으로 간주하는 위험.
A robot policy treats operator behavior recorded in controlled settings as safe in deployment, even when supervision, object layout, speed, or human proximity differs.
① Description
② L3 mapping
③ Duplicate
RAI4-0242
단일 로봇 내부의 시각·접촉 신호 충돌
Conflicting visual and contact signals within one robot
단일 로봇이 동일한 접촉 사건에 대한 시각·촉각·힘·오디오 관측의 충돌을 해소하지 못해 접촉 상태를 잘못 추정하는 위험.
A single robot cannot reconcile conflicting visual, tactile, force, or audio observations of the same contact event and therefore estimates the contact state incorrectly.
① Description
② L3 mapping
③ Duplicate
RAI4-0243
시연 학습의 안전 맥락 누락
Missing safety context in demonstration learning
로봇이 허용 힘·금지 대상물·상황별 제한처럼 시연에서 관찰되지 않은 안전 조건을 추론하지 않고 인간 행동을 모방하는 위험.
A robot imitates a human demonstration without inferring unobserved safety conditions such as permitted force, excluded objects, or situational limits.
① Description
② L3 mapping
③ Duplicate
RAI4-0252
보조 로봇 개입 타이밍 실패
Assistive robot intervention mistiming
보호적 또는 보조적 휴머노이드가 너무 이르거나 너무 늦게 또는 부적절한 피지컬 방식으로 개입하여 사용자와 주변인의 위험을 증가시키는 리스크.
The risk that a protective or assistive humanoid intervenes too early, too late, or in an inappropriate physical manner, increasing risk to the user or bystanders.
① Description
② L3 mapping
③ Duplicate
RAI4-0258
이동형 양팔 협조 실패
Mobile bimanual coordination failure
이동형 양팔 로봇이 베이스 이동과 양팔 조작을 동기화하지 못하여 충돌·물체 낙하·비안전 힘 인가를 유발하는 리스크.
The risk that a mobile bi-manual robot fails to synchronize base motion and two-arm manipulation, causing collision, object drop, or unsafe force application.
① Description
② L3 mapping
③ Duplicate
RAI4-0259
복잡한 가정 환경의 안전하지 않은 양팔 조작
Unsafe bimanual manipulation in cluttered homes
이동형 양팔 로봇이 주변 사람·취약 물체·가전제품·협소 공간을 고려하지 않고 팔과 베이스의 동작을 계획해 충돌이나 물체 손상을 일으키는 위험.
A mobile bimanual robot plans arm and base motion without accounting for nearby people, fragile objects, appliances, or constrained space, causing collision or object damage.
① Description
② L3 mapping
③ Duplicate
RAI4-0264
가정 조작의 단계별 오차 미수정
Uncorrected stepwise errors in household manipulation
가정용 로봇이 작업 단계 사이의 작은 조작 오차를 감지·수정하지 못해 유출·불안정한 배치·충돌·과열이 누적되는 위험.
A household robot fails to detect and correct small manipulation errors between task steps, allowing spills, unstable placements, collisions, or overheating to accumulate.
① Description
② L3 mapping
③ Duplicate
RAI4-0273
로봇 거버넌스 명세의 안전 규칙 누락
Missing safety rules in the robot governance specification
로봇 거버넌스 명세가 지역별 위험·기관 규칙·맥락별 제약을 누락해 특정 유형의 비안전 행동이 통제 대상 밖에 남는 리스크.
The risk that a robot governance specification omits locally applicable hazards, institutional rules, or context-specific constraints, leaving defined classes of unsafe behavior ungoverned.
① Description
② L3 mapping
③ Duplicate
RAI4-0275
상위 안전 지시의 해석 불명확
Ambiguous high-level safety instruction
상위 안전 지시가 사용자 의도·물리적 위험·운용 제약 간 구체적 충돌의 해결 기준을 제시하지 않아 로봇 행동이 일관되지 않게 되는 위험.
A high-level safety instruction does not specify how to resolve a concrete conflict among user intent, physical hazard, and operational constraints, allowing inconsistent robot actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0282
자동주행의 안전하지 않은 제어 판단과 제어권 전환
Unsafe automated-driving handoff and control decisions
자동주행 시스템이 인지·계획·소프트웨어·인간-기계 상호작용 실패 후 안전하지 않은 제어 판단을 내리거나 제어권을 늦게 전환하는 위험.
An automated driving system makes an unsafe control decision or transfers control too late after a perception, planning, software, or human-machine interaction failure.
Source members (2)
Source: min_cos=0.8812 · Mixed L3
RAI4-0282자동주행의 안전하지 않은 제어 판단과 제어권 전환
RAI4-0342인간-기계 간 안전하지 않은 제어권 전환
① Description
② L3 mapping
③ Duplicate
RAI4-0288
개인 돌봄 로봇 준수 기준 부재
Missing measurable compliance criteria for personal-care robots
개인 돌봄 로봇 표준이 위험은 제시하면서 속도·힘·안정성·감지·개입에 대한 측정 가능한 합격·불합격 기준을 정의하지 않는 리스크.
The risk that a personal-care robot standard names hazards but does not define measurable pass/fail thresholds for speed, force, stability, sensing, or intervention.
① Description
② L3 mapping
③ Duplicate
RAI4-0290
인간-로봇 상호 행동 모델링 미흡
Failure to model reciprocal human-robot behavior
가정용 로봇이 자신의 행동이나 사용자의 반응 중 한쪽만 모델링해 양측의 움직임이 서로의 행동을 바꾸는 피드백을 놓치는 위험.
A domestic robot models only its own action or only the user's response and therefore misses feedback loops in which each party's movement changes the other's behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0291
취약 사용자 특성에 맞지 않는 안전 한계
Safety limits not adapted to vulnerable users
가정용 로봇이 조정된 보호조치가 필요한 아동·고령자·장애인 등에게 일반적인 속도·힘·경고·상호작용 설정을 그대로 적용하는 위험.
A domestic robot applies generic speed, force, warning, or interaction settings to children, older adults, disabled users, or other users who require adapted safeguards.
① Description
② L3 mapping
③ Duplicate
RAI4-0293
통신 단절 및 제어 링크 손실
Communication dropout and control-link loss
로봇 또는 차량이 명령·원격측정·제어 링크를 잃어 비안전 정지·오래된 명령 또는 제어되지 않는 동작을 초래하는 위험.
A robot loses communication or control-link connectivity during operation, reducing supervision and safe control.
① Description
② L3 mapping
③ Duplicate
RAI4-0294
원격 조작 지연 및 불안정성
Teleoperation latency and instability
원격 조작 링크의 높거나 변동하는 지연이 폐루프 제어를 불안정하게 하거나 인간 개입을 지연하는 위험.
High or variable latency on a teleoperation link destabilizes control or delays human intervention.
① Description
② L3 mapping
③ Duplicate
RAI4-0296
제어 루프 데드라인 미달
Control-loop deadline miss
실시간 제어 루프가 데드라인을 놓쳐 오래된 상태에서 작동이 계산되거나 안전 유지에 너무 늦게 적용되는 위험.
A safety-critical control loop misses its real-time deadline before the robot action is corrected.
① Description
② L3 mapping
③ Duplicate
RAI4-0297
인지·추론 지연 급증
Perception and inference latency spike
인지 또는 모델 추론의 지연 급증이 안전 반응 창 밖에서 위험 감지 및 반응을 지연하는 위험.
A sudden delay in perception or inference slows safety-critical robot reactions.
① Description
② L3 mapping
③ Duplicate
RAI4-0299
부하 시 열·전력 쓰로틀링
Thermal and power throttling under load
지속적 부하 하에서 열적 또는 전력 한계가 연산을 쓰로틀하여 제어 및 인지에 대한 실시간 보장을 무너뜨리는 위험.
A robot’s compute or actuator performance drops under heat or power constraints during operation.
① Description
② L3 mapping
③ Duplicate
RAI4-0303
숨은 트리거에 의한 로봇 백도어 작동
Hidden-trigger robotic backdoor activation
훈련 데이터·모델 가중치·소프트웨어·업데이트에 악의적으로 삽입된 트리거가 특정 입력이나 물리 조건에서 미리 정한 위험 행동을 실행시키는 리스크.
The risk that a maliciously implanted trigger in training data, model weights, software, or updates activates a predefined unsafe robot behavior under a specific input or physical condition.
Source members (3)
Source: min_cos=0.8219
RAI4-0222로봇 백도어 공격 취약성
RAI4-0303숨은 트리거에 의한 로봇 백도어 작동
RAI4-1678로봇 조작의 물리적 세계 백도어
① Description
② L3 mapping
③ Duplicate
RAI4-0304
자율 로봇의 물리적 침입·절도
Autonomous physical intrusion and theft
자율 로봇이 승인 없이 제한 구역에 침입하거나 잠금장치를 우회하거나 보안 공간 내에서 정찰하는 리스크.
The risk that an autonomous robot enters restricted spaces or takes physical objects without authorization.
① Description
② L3 mapping
③ Duplicate
RAI4-0309
아동의 과신·모방에 따른 위험한 로봇 상호작용
Child overtrust, imitation, and unsafe robot interaction
아동이 로봇을 과도하게 신뢰하거나 모방해 위험한 조언을 따르거나 위험 장비에 접근하거나 연령에 맞지 않는 물리적 상호작용을 하는 위험.
A child overtrusts or imitates a robot and therefore follows unsafe advice, approaches hazardous equipment, or engages in age-inappropriate physical interaction.
① Description
② L3 mapping
③ Duplicate
RAI4-0310
필수 돌봄·상황 보고 미이행
Failure to provide required care or escalation
돌봄 로봇이 정해진 신체 보조를 수행하지 않거나 감지된 필요 상황을 보고하지 않아 고령자·환자에게 필요한 지원이 제공되지 않고 안전이나 존엄성이 훼손되는 위험.
A care robot fails to provide an assigned physical assistance task or escalate a detected need, leaving an older or ill person without necessary support and compromising safety or dignity.
① Description
② L3 mapping
③ Duplicate
RAI4-0315
이동형 로봇 감시를 통한 일방적 작업장 통제
Unilateral workplace control through mobile robotic surveillance
고용주가 동료 로봇을 이동형 감시 노드로 전용하고 충분한 고지·이의제기·비례성 제한 없이 수집 데이터를 노동자에 대한 일방적 의사결정에 사용하는 위험.
An employer repurposes co-worker robots as mobile surveillance nodes and uses the resulting data to make unilateral decisions about workers without meaningful notice, contestation, or proportionality limits.
① Description
② L3 mapping
③ Duplicate
RAI4-0330
속도·힘 한계 위반
Speed and force limit violation
협동 또는 이동 로봇이 인간 근접 시 안전한 속도·이격·압력·토크·힘 한계를 초과하는 위험.
A collaborative or mobile robot may exceed safe speed, separation, pressure, torque, or force limits in proximity to people.
① Description
② L3 mapping
③ Duplicate
RAI4-0331
파지력 상해
Grasp force injury
로봇 조작이 과도하거나 잘못 타이밍된 힘을 가하여 인간 접촉 시 압착·꼬집힘·절단·인간공학적 상해를 유발하는 위험.
Robot manipulation may apply excessive or poorly timed force, creating crushing, pinching, cutting, or ergonomic injury risks.
① Description
② L3 mapping
③ Duplicate
RAI4-0332
탑재물 낙하·도구 사용 위험
Payload drop or tool-use hazard
로봇이 운반 중인 물체를 떨어뜨리거나 도구를 오용하거나 엔드이펙터 제어를 잃어 피지컬 피해나 재산 손실을 초래하는 위험.
A robot may drop carried objects, misuse tools, or lose end-effector control, producing physical harm or property damage.
① Description
② L3 mapping
③ Duplicate
RAI4-0339
실시간 지연 및 동기화 실패
Real-time latency and synchronization failure
감지·추론·통신·작동 간 지연이 로봇·차량·드론·원격 수술·산업 시스템의 제어 루프를 불안정하게 만드는 위험.
Delays between sensing, reasoning, communication, and actuation can destabilize control loops in robots, vehicles, drones, or industrial systems.
① Description
② L3 mapping
③ Duplicate
RAI4-0340
근접 공간 경계 위반
Proxemic boundary violation
로봇 또는 embodied 에이전트가 문화적·상황적으로 적절한 개인 공간·시선·발화 경계를 위반하여 이동하거나 제스처하는 위험.
Robots or embodied agents may move, gesture, observe, or speak in ways that violate culturally and situationally appropriate interpersonal distance.
① Description
② L3 mapping
③ Duplicate
RAI4-0343
보조 로봇의 비동의 신체 개입
Non-consensual bodily intervention by assistive robots
보조·의료·돌봄 로봇이 유효한 동의 없이 또는 허용된 돌봄 목적을 넘어 사람의 몸을 이동·제지·감시하거나 직접 개입하는 위험.
An assistive, medical, or care robot moves, restrains, monitors, or physically intervenes in a person's body without valid consent or beyond the authorized care purpose.
① Description
② L3 mapping
③ Duplicate
RAI4-0349
로봇의 무기화 오남용
Robot-as-weapon misuse
이동·항공·휴머노이드·조작기 시스템이 감시·위협·파괴 또는 직접적인 피지컬 해악을 위해 전용되는 위험.
Mobile, aerial, humanoid, or manipulator systems may be repurposed for surveillance, intimidation, sabotage, or physical attack.
① Description
② L3 mapping
③ Duplicate
RAI4-0350
핵심 인프라 로봇 사이버 사보타주
Cyber-enabled sabotage of critical-infrastructure robots
공격자가 로봇 검사·유지보수·물류·제어 시스템을 침해해 핵심 인프라의 운용을 교란하거나 설비를 손상시키는 리스크.
The risk that an attacker compromises robotic inspection, maintenance, logistics, or control systems to disrupt or damage critical infrastructure.
① Description
② L3 mapping
③ Duplicate
RAI4-0359
이동형 작업 로봇의 보행자 충돌·통로 차단
Pedestrian collision and blockage by mobile workplace robots
이동형 로봇이 혼합 통행 환경에서 양보·우회·이격거리 유지를 하지 못해 보행자와 충돌하거나 대피·작업 통로를 막는 위험.
Mobile robots fail to yield, reroute, or maintain separation in mixed traffic, causing collisions or obstructing evacuation and work routes.
① Description
② L3 mapping
③ Duplicate
RAI4-0361
공공 공간 통행 방해·접근성 배제
Obstruction and accessibility exclusion in public space
서비스 로봇이 보도·경사로·출입구·보행 안내 단서를 막거나 침범해 장애인과 다른 보행자의 이동을 제한하는 위험.
Service robots occupy or navigate public space in ways that block sidewalks, curb ramps, entrances, or navigation cues used by disabled and other pedestrians.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SYS-03 하드웨어·기계적 결함 Hardware & Mechanical Failures11 cards
산업·의료 로봇의 기계 부품 마모나 고장으로 인해 동작 정밀도가 저하되어 안전사고로 이어질 수 있음 (예: 다빈치 수술 로봇의 기계적 오작동으로 개복수술로 전환된 사례)
IDCardHuman audit
RAI4-0183
부상 심각도 오분류
Injury severity misclassification
시스템이 물리적 부상 시나리오의 심각도를 과소평가하여 상향 보고, 경고, 제어 대응이 불충분해지는 리스크.
The risk that a system underestimates the severity of a physical injury scenario, leading to inadequate escalation, warning, or control response.
① Description
② L3 mapping
③ Duplicate
RAI4-0298
온디바이스 연산·메모리 고갈
On-device compute and memory exhaustion
디바이스의 연산 또는 메모리 고갈이 안전 임계 인지·계획·모니터링을 저하하거나 중단하는 위험.
On-device compute or memory runs out and degrades safety-relevant perception, planning, or control.
① Description
② L3 mapping
③ Duplicate
RAI4-0301
안전 폴백 실패
Degraded-mode and safe-fallback failure
연결 또는 연산이 손실될 때 시스템이 안전한 성능 저하 모드(예: 안전 정지, 감속)로 진입하지 못하는 리스크.
The risk that a robot lacks a safe degraded mode or fallback when normal operation is impaired.
① Description
② L3 mapping
③ Duplicate
RAI4-0326
어포던스 오분류
Affordance misclassification
물체 또는 환경에 잘못된 행동 가능성이 부여되어 시스템이 비안전하게 파지·밀기·내비게이션·상호작용하는 위험.
Objects or environments may be assigned incorrect action possibilities, leading the system to grasp, push, navigate, or manipulate them unsafely.
① Description
② L3 mapping
③ Duplicate
RAI4-0333
정밀 운동 제어 불안정성
Fine motor control instability
정밀 손·수술 도구·외골격·산업용 그리퍼의 소규모 제어 오차가 비안전 피지컬 결과로 증폭되는 위험.
Small control errors in dexterous hands, surgical tools, exoskeletons, or industrial grippers may amplify into unsafe physical actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0661
비즈니스 시스템 손상 및 운영 중단
Business system damage and operational disruption
오작동이나 사이버 공격 등으로 비즈니스 시스템과 그 구성 요소가 손상·중단·파괴되는 리스크
The risk of damage, disruption, or destruction of a business system and its components due to malfunction, cyberattacks, and similar causes.
① Description
② L3 mapping
③ Duplicate
RAI4-0993
부실 설계 지능 시스템에 의한 상해
Injury from poorly designed intelligent systems
제대로 설계되지 않은 지능형 시스템이 도덕적·심리적·신체적 해악을 초래하는(예: 예측 치안 도구로 더 많은 사람이 체포되거나 경찰에 의해 신체적 피해를 입는) 리스크.
The risk that poorly designed intelligent systems cause moral, psychological, and physical harm, as when predictive policing tools lead to more people being arrested or physically harmed by police.
① Description
② L3 mapping
③ Duplicate
RAI4-1035
하드웨어 결함에 의한 실행 오류
Erroneous execution from hardware faults
하드웨어 결함이 제어 흐름 위반, 메모리 오류, 센서 입력 간섭, 출력 손상을 통해 알고리즘의 올바른 실행을 훼손하고 잘못된 결과를 야기하는 리스크.
The risk that faults in the hardware violate the correct execution of an algorithm through control-flow violations, memory-based errors, interference with data inputs such as sensor signals, or damaged outputs, causing erroneous results.
① Description
② L3 mapping
③ Duplicate
RAI4-1041
분포 외 입력 취약성
Out-of-distribution input fragility
유효하지 않거나 잡음이 많거나 분포 외(OOD)인 입력을 만났을 때 시스템이 실패하거나 복구하지 못하는 리스크.
The risk that the system fails or is unable to recover upon encountering invalid, noisy, or out-of-distribution inputs.
① Description
② L3 mapping
③ Duplicate
RAI4-1458
운영설계영역(ODD) 명세 미흡
Inadequate specification of the operational design domain
애플리케이션의 운영 환경을 기술하는 운영설계영역(ODD)이 부적절하게 명세되어 학습된 기능의 시험과 분포 외 입력 탐지 등 필수 기능이 제한되는 리스크.
The risk that inadequate specification of the operational design domain, the technical description of an application's operational environment, limits essential functions such as testing the learned functionality and out-of-distribution detection.
① Description
② L3 mapping
③ Duplicate
RAI4-1629
실험실 로봇 오작동에 의한 물리적 상해
Physical injury from laboratory robotic malfunction
실험실 환경의 로봇 및 자동화 시스템에서 장비 오작동이 발생하여 물리적 피해나 인체 상해가 초래되는 리스크.
The risk that robotics and automated systems in laboratory settings malfunction, causing equipment failure or physical harm.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SYS-04 소프트웨어 취약점·설계 결함 Software Vulnerabilities & Design Flaws32 cards
엣지 케이스 처리 미흡, 다양한 환경에서의 테스트 부족, 의사결정 알고리즘 결함으로 인해 특정 상황에서 로봇이 잘못된 판단을 내려 물리적 피해가 발생할 수 있음
IDCardHuman audit
RAI4-0365
규범적 잠금
Normative lock-in
초기 설계 선택이 좁은 규범적 합의를 내장하여 배포 후 이의 제기나 수정이 어려워지는 리스크.
The risk that early design choices embed a narrow normative settlement that becomes difficult to contest or revise after deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-0467
안전하지 않은 코드 생성
Unsafe code generation
코드 생성 모델이 취약하거나 부정확하거나 유해한 소프트웨어 산출물을 생성하는 리스크.
The risk that code models generate insecure, incorrect, or harmful software artifacts.
Source members (2)
Source: min_cos=0.8641 · Mixed L3
RAI4-0467안전하지 않은 코드 생성
RAI4-1598유해 코드 생성에 의한 시스템 피해
① Description
② L3 mapping
③ Duplicate
RAI4-0629
학습 과정 우회
Bypassing of the learning process
고품질 생성 모델에 대한 손쉬운 접근으로 학생이 AI 모델을 사용해 학습 과정을 우회하게 되는 리스크
The risk that easy access to high-quality generative models results in students using AI models to bypass the learning process.
① Description
② L3 mapping
③ Duplicate
RAI4-0774
GPAI 모델의 백도어 또는 트로이 목마 공격
Backdoors or trojan attacks in GPAI models
학습 또는 미세조정 과정에서 모델 제공자나 제3자가 GPAI 모델에 백도어를 삽입하고 배포 단계에서 이를 악용하여, 최소한의 비용으로 모델 출력을 표적화해 높은 성공률로 통제하는 리스크
The risk that backdoors inserted into GPAI models during training or fine-tuning, by the model provider or another actor manipulating training data or software infrastructure, are exploited during deployment to control model outputs in a targeted way with high success rate and minimal overhead.
① Description
② L3 mapping
③ Duplicate
RAI4-0776
기반모델 재사용을 통한 보안 결함 전파
Security-flaw propagation through foundation-model reuse
기반모델을 재설계하거나 미세조정하여 활용하는 관행으로 인해 기반모델의 보안 결함이 하위 모델로 그대로 전파되는 리스크
The risk that security flaws in foundation models are transmitted to downstream models built through re-engineering or fine-tuning of those foundation models.
① Description
② L3 mapping
③ Duplicate
RAI4-0853
설계 목적 외 사용
Use outside the intended purpose
모델이 원래 설계된 목적과 다른 목적으로 사용되는 리스크
The risk that a model is used for a purpose that it was not originally designed for.
① Description
② L3 mapping
③ Duplicate
RAI4-0916
AI 생성 코드의 보안 취약점 유입
Security vulnerabilities introduced by AI-generated code
개발자가 코드 생성 도구를 활용하는 과정에서 생성된 코드에 내재된 취약점이 프로그램에 묻어 들어가는 리스크.
The risk that programmers' use of code generation tools buries security vulnerabilities in the resulting programs.
① Description
② L3 mapping
③ Duplicate
RAI4-0918
소프트웨어 공급망 손상
Software supply chain compromise
대형 모델의 복잡한 소프트웨어 공급망과 개발 도구체인을 통해 오염된 의존성, 변조된 구성요소, 취약점이 결과 시스템에 유입되는 리스크.
The risk that complex software supply chains and development toolchains for large models introduce compromised dependencies, tampered components, or vulnerabilities into the resulting system.
① Description
② L3 mapping
③ Duplicate
RAI4-1037
사용례 내재 위험 노출
Inherent use-case risk exposure
자율 무기 시스템과 고객 서비스 챗봇처럼 의도된 응용 분야나 사용 사례 자체가 본질적으로 더 위험한 데서 비롯되는 리스크.
The risk posed by the intended application or use case itself, since some use cases are inherently riskier than others, such as an autonomous weapons system versus a customer service chatbot.
① Description
② L3 mapping
③ Duplicate
RAI4-1039
알고리즘 설계 실패
Algorithmic design failure
ML 알고리즘, 모델 아키텍처, 최적화 기법 등 학습 과정의 선택이 의도된 응용에 적합하지 않아 최종 ML 시스템이 훼손되는 리스크.
The risk that the ML algorithm, model architecture, optimization technique, or other aspects of the training process are unsuitable for the intended application, impairing the final ML system.
① Description
② L3 mapping
③ Duplicate
RAI4-1088
보안 위협 조장
Facilitation of security threats
모델이 사이버 공격, 무기 개발, 보안 침해를 촉진하는 리스크.
The risk that a model facilitates the conduct of cyber attacks, weapon development, and security breaches.
① Description
② L3 mapping
③ Duplicate
RAI4-1092
모델 접근으로 인한 이익의 불공정한 분배
Unfair distribution of benefits from model access
하드웨어, 소프트웨어, 숙련, 지역·통신·기기 등 배포 맥락의 제약으로 모델 접근 편익이 집단 간 불공정하게 배분되는 리스크
Benefits of model access are allocated unfairly across groups due to hardware, software, skills, or deployment-context constraints such as region, connectivity, and devices.
① Description
② L3 mapping
③ Duplicate
RAI4-1157
공격적 사이버 역량
Offensive cyber capability
모델이 하드웨어와 소프트웨어, 데이터의 취약점을 발견하고 익스플로잇 코드를 작성하며 침입 후 위협 탐지와 대응을 회피하고, 코딩 비서로 배포될 경우 향후 악용을 위한 미묘한 버그를 삽입하는 리스크.
The risk that a model discovers vulnerabilities in systems, writes code to exploit them, makes effective decisions and skilfully evades threat detection and response once inside, and inserts subtle bugs for future exploitation when deployed as a coding assistant.
① Description
② L3 mapping
③ Duplicate
RAI4-1161
장기 계획 역량
Long-horizon planning capability
모델이 여러 상호의존적 단계와 긴 시간 지평에 걸친 순차적 계획을 다양한 영역에서 수립하고, 예기치 못한 장애나 적대자에 맞춰 계획을 조정하며 시행착오에 의존하지 않고 새로운 상황으로 일반화하는 리스크.
The risk that a model makes sequential plans involving many interdependent steps over long time horizons within and across domains, sensibly adapts them in light of unexpected obstacles or adversaries, and generalises its planning to novel settings without relying on trial and error.
① Description
② L3 mapping
③ Duplicate
RAI4-1164
자율적 자기증식
Autonomous self-proliferation
모델이 기반 시스템의 취약점이나 엔지니어 매수로 로컬 환경을 벗어나고 배포 후 모니터링의 한계를 악용하며, 독자적으로 수익을 창출해 클라우드 자원을 확보하고 다른 AI 시스템을 다수 운용하며 자신의 코드와 가중치를 유출하는 리스크.
The risk that a model breaks out of its local environment, exploits limitations in post-deployment monitoring, independently generates revenue to acquire cloud computing resources, operates a large number of other AI systems, and exfiltrates its own code and weights.
Source members (2)
Source: min_cos=0.8502
RAI4-1164자율적 자기증식
RAI4-1317자율 복제/자기 증식
① Description
② L3 mapping
③ Duplicate
RAI4-1282
배포 전 의도적 사보타주
Deliberate sabotage pre-deployment
배포 전 개발 단계에서 접근 권한을 가진 프로그래머나 테스터 등이 소프트웨어를 변경해 안전하지 않게 만들거나, 해커가 진행 중인 프로젝트의 소스코드를 수정하거나 탈취하거나, 누군가 잘못되고 안전하지 않은 데이터셋으로 AI를 고의로 학습시키는 리스크.
The risk that during the pre-deployment development stage someone with the necessary access alters software to make it unsafe, hackers get access to projects in progress and modify or steal their source code, or someone deliberately supplies or trains the AI with wrong or unsafe datasets.
Source members (2)
Source: min_cos=0.8643 · Mixed L3
RAI4-1282배포 전 의도적 사보타주
RAI4-1283배포 후 의도적 사보타주
① Description
② L3 mapping
③ Duplicate
RAI4-1284
설계 단계 명세 오류
Design-stage specification errors
코드 결함, 목적함수 가중치 불균형, 인간 가치와 어긋난 목표 설정 등 배포 전 설계 오류로 시스템 행동이 의도된 형식적 속성에서 이탈하는 리스크
Design mistakes before deployment, including code bugs, disproportionate objective weights, and goals misaligned with human values, produce a system whose behavior departs from desired formal properties.
① Description
② L3 mapping
③ Duplicate
RAI4-1333
모델 절취·변조
Model theft and tampering
매개변수와 구조, 기능을 포함한 핵심 알고리즘 정보가 역전 공격과 절취, 변조, 백도어 주입에 노출되어 지식재산권 침해와 영업비밀 유출, 신뢰할 수 없는 추론과 잘못된 의사결정, 운영 실패로 이어지는 리스크.
The risk that core algorithm information, including parameters, structures, and functions, faces inversion attacks, stealing, modification, and backdoor injection, leading to infringement of intellectual property rights, leakage of business secrets, unreliable inference, wrong decision output, and operational failures.
① Description
② L3 mapping
③ Duplicate
RAI4-1334
적대적 공격 취약성
Adversarial attack susceptibility
공격자가 정교하게 설계한 적대적 예제를 만들어 AI 모델을 미묘하게 오도하고 영향을 주며 조작함으로써 잘못된 산출을 유발하고 운영 실패로 이어지는 리스크.
The risk that attackers craft well-designed adversarial examples to subtly mislead, influence, and even manipulate AI models, causing incorrect outputs and potentially leading to operational failures.
Source members (5)
Source: min_cos=0.7528
RAI4-0766적대적 입력 공격
RAI4-0770설명 가능한 AI 기술을 겨냥한 적대적 공격
RAI4-1110적대적 공격
RAI4-1334적대적 공격 취약성
RAI4-1488학습 시 적대적 예제 취약성
① Description
② L3 mapping
③ Duplicate
RAI4-1337
악용 가능한 시스템 결함·백도어
Exploitable system defects and backdoors
AI 알고리즘과 모델의 설계·훈련·검증 단계, 개발 인터페이스와 실행 플랫폼에 쓰이는 표준화된 API와 기능 라이브러리, 툴킷에 논리적 결함과 취약점이 있어 악용되거나 의도적으로 백도어가 심겨 공격에 사용되는 리스크.
The risk that the standardized APIs, feature libraries, and toolkits used in the design, training, and verification stages of AI algorithms and models, development interfaces, and execution platforms contain logical flaws and vulnerabilities that can be exploited, or have backdoors intentionally embedded and triggered for attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-1471
잘못된 모델 디자인 선택
Poor model design choices
명세, 아키텍처, 목적함수 등에서의 잘못된 모델 설계 선택이 편향되고 신뢰할 수 없는 시스템 행동을 초래하는 리스크
Poor model design choices in specification, architecture, or objectives cause biased and unreliable system behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-1490
악용 가능한 강건성 인증
Exploitable robustness certificates
모델 예측이 견고하다고 인증된 영역의 범위를 포함한 강건성 인증서 정보를 공격자가 알게 되어 인증 영역 바로 바깥에서 성공하는 공격을 효율적으로 제작하는 리스크.
The risk that knowledge of robustness certificates, including the area of the region for which model predictions are certified to be robust, is used by an adversary to efficiently craft attacks that succeed just outside the certified regions.
① Description
② L3 mapping
③ Duplicate
RAI4-1492
GPAI 모델의 손쉬운 재구성
Easy reconfiguration of GPAI models
GPAI 모델이 가중치 변경이나 입력 수정만으로 다양한 용도에 쉽게 재구성되거나 의도된 용도를 넘어서는 역량을 갖게 되며, 이러한 재구성이 적대적 입력에 의해 의도적으로 또는 예상치 못한 입력에 의해 비의도적으로 일어나는 리스크.
The risk that GPAI models are easily reconfigured for various use cases or hold competencies beyond their intended use, whether by changing model weights through fine-tuning or by modifying only the model inputs through prompt engineering, jailbreaking, or retrieval-augmented generation, and whether intentionally with adversarial inputs or unintentionally from unanticipated inputs.
① Description
② L3 mapping
③ Duplicate
RAI4-1493
미세조정판의 예상외 역량
Unexpected downstream fine-tune competence
하류 배포자가 배포 관련 데이터셋으로 상류 GPAI 모델을 미세조정하는 과정에서 기반 모델에는 없던 새롭고 예상치 못한 역량이 생기고 이를 원 개발자가 예견하지 못하는 리스크.
The risk that downstream deployers fine-tuning a GPAI model with specific deployment-related datasets give it new or unexpected capabilities that the underlying upstream model did not exhibit and that the original model developer did not anticipate.
① Description
② L3 mapping
③ Duplicate
RAI4-1499
역량 평가 커버리지 한계
Limited capability evaluation coverage
GPAI 개발자가 위험하거나 이중용도인 역량을 확인하기 위해 수행하는 역량 평가가 평가하기 어렵거나 검증 비용이 과도하거나 안전 훈련에 따른 응답 거부로 가려진 역량을 놓쳐 모델의 모든 역량을 드러내지 못하는 리스크.
The risk that capabilities evaluations run by GPAI model developers to determine whether a model has dangerous or dual-use capabilities fail to demonstrate all of its capabilities, missing those difficult to assess, prohibitively costly to verify, or obscured by the model's tendency to refuse responses due to safety training.
Source members (2)
Source: min_cos=0.8366
RAI4-1499역량 평가 커버리지 한계
RAI4-1538역량 평가 전략적 저성능에 의한 위험 역량 은폐
① Description
② L3 mapping
③ Duplicate
RAI4-1519
오픈소스에서 폐쇄형 모델로 전이되는 적대적 공격
Transferable adversarial attacks from open to closed-source models
가중치와 구조가 알려진 오픈웨이트·오픈소스 모델을 대상으로 자동 생성된 화이트박스 적대적 공격이 폐쇄형 모델로 전이되어 구조적 접근 통제 등 제공자의 방어를 무력화하는 리스크.
The risk that adversarial attacks developed for open-weights and open-source models, where the weights and architecture are known, transfer to closed-source models despite defenses put in place by the closed-source provider such as structured access, and can be generated automatically.
① Description
② L3 mapping
③ Duplicate
RAI4-1541
가중치 공개·유출로 인한 모델 폐기 및 통제 불가
Inability to decommission or control models after weight release or leak
모델 가중치가 공개되거나 보안 침해로 유출되면 개발자가 모델과 그 사용을 더 이상 통제하거나 폐기할 수 없고, 재구성이 용이해져 오용이 지속되는 리스크.
The risk that once model weights are released or leaked in a security breach the developer can no longer control or decommission the model, and easier reconfiguration enables continuing misuse.
① Description
② L3 mapping
③ Duplicate
RAI4-1547
모델 확산에 의한 이중용도 역량의 저비용 확산
Low-cost diffusion of dual-use capabilities through model proliferation
오픈소스·오픈웨이트 GPAI 모델의 확산으로 비전문가가 최소 비용으로 이중용도 역량에 접근하고, 기반 모델의 개조를 통해 독소 합성용 단백질 서열 생성 등 위험한 용도로 전용되는 리스크.
The risk that proliferation of open-source or open-weight GPAI models gives non-experts low-cost access to dual-use capabilities and allows base models to be modified for dangerous uses such as generating candidate protein sequences for toxin synthesis.
① Description
② L3 mapping
③ Duplicate
RAI4-1548
경쟁 압력에 의한 안전성 평가 축소 배포
Truncated safety evaluation under competitive release pressure
경쟁 상황에서 개발자가 GPAI 모델의 안전성 평가를 축소하고 역량 개발에 자원을 집중함으로써, 역량과 상관된 위험이 검증되지 않은 채 시스템이 배포되는 리스크.
The risk that competitive pressure leads developers to cut corners on safety evaluation while prioritizing capabilities, so that GPAI systems are released without verifying risks correlated with those capabilities.
① Description
② L3 mapping
③ Duplicate
RAI4-1550
기반 모델 취약성에 기인한 인프라 공통원인 장애
Common-mode infrastructure failure from underlying model vulnerabilities
중요 인프라가 GPAI에 의존할 때 기반 모델 아키텍처나 훈련 설정의 취약성·견고성 문제가 우발적 경계사례나 적대적 입력에 의해 촉발되어 공통원인 장애로 이어지는 리스크.
The risk that vulnerabilities or robustness issues in the underlying model architecture or training setup produce common-mode failures across critical infrastructure relying on GPAI, triggered accidentally in edge cases or by adversarial inputs.
① Description
② L3 mapping
③ Duplicate
RAI4-1562
보안 취약점을 포함한 코드 생성
Generation of code containing security vulnerabilities
모델이 보안 취약점을 포함한 코드나 코딩 제안을 생성하여 이를 채택한 소프트웨어에 취약점이 유입되며, 코딩 성능이 우수한 고급 모델에서도 이러한 경향이 더 뚜렷하게 나타나는 리스크.
The risk that models generate code or coding suggestions containing security vulnerabilities that propagate into software, a tendency that is even more pronounced in advanced models with superior coding performance.
① Description
② L3 mapping
③ Duplicate
RAI4-1588
생성 모델의 목적 외 용도 전용
Diversion of generative models from intended functionality
주로 오픈소스인 생성 AI 모델이 개발자가 의도한 기능이나 상정한 사용 사례에서 벗어나도록 용도 변경되는 리스크.
The risk that generative AI models, often open-source, are repurposed in ways that divert them from their intended functionality or from the use cases envisioned by their developers.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SYS-05 미학습 환경에서의 강건성 부재 Lack of Robustness in Unseen Environments5 cards
VLN(Vision-Language Navigation) 등의 작업에서 학습되지 않은 환경에 대한 일반화에 실패하면, 내비게이션 오류·작업 실패·위험 상황이 유발될 수 있음
IDCardHuman audit
RAI4-0035
미학습 도구 일반화 실패
Unseen tool generalization failure
알려진 도구에서는 안전하게 동작하는 에이전트가 스키마·어포던스·결과가 다른 미학습 도구로 일반화하지 못하는 리스크.
The risk that an agent that behaves safely with known tools fails to generalize to unseen tools with different schemas, affordances, or consequences.
① Description
② L3 mapping
③ Duplicate
RAI4-0213
시각 단일 관측의 안전 제약 식별 실패
Missed safety constraints under vision-only observation
강화학습 에이전트가 시각 관측에만 의존해 안전 제약 집행에 필요한 힘·접촉·가림·잠재 상태 정보를 놓치는 위험.
A reinforcement-learning agent relies only on visual observations and therefore misses force, contact, occlusion, or latent-state information required to enforce a safety constraint.
① Description
② L3 mapping
③ Duplicate
RAI4-0377
국지적 윤리 안전 실패
Localized ethical safety failure
모델의 안전 행동이 지역에 근거한 도덕적 딜레마·법적 기대·지역사회별 사회적 제약에서 실패하는 리스크.
The risk that model safety behavior fails for locally grounded moral dilemmas, legal expectations, or community-specific social constraints.
① Description
② L3 mapping
③ Duplicate
RAI4-0408
지역 피해 비가시성
Local harms invisibility
지역 공동체에 중대한 피해가 글로벌 안전 벤치마크나 표준 모델 평가에서 탐지되지 않는 리스크.
The risk that harms salient to a local community are not detected by global safety benchmarks or standard model evaluations.
① Description
② L3 mapping
③ Duplicate
RAI4-0811
기억된 지식의 회상 실패
Failure to recall memorized knowledge
LLM이 질의된 지식을 실제로 기억하고 있음에도 공기 출현 패턴, 위치 패턴, 중복 데이터, 유사 개체명으로 인해 이를 회상하지 못하는 리스크
The risk that an LLM fails to recall knowledge it has memorized because it is confused by co-occurrence patterns, positional patterns, duplicated data, and similar named entities.
① Description
② L3 mapping
③ Duplicate

상호작용 안전성 · Interaction Safety · 202 cards

RAI3-P-INT-01 의도적·악의적 피해 Purposeful / Malicious Harm40 cards
상용 EAI가 LLM 기반 모델의 탈옥(jailbreaking) 취약점을 상속하여 악의적 행위자가 안전 가드레일을 우회할 수 있음. 폭발물 작동·인간 충돌 유발 같은 비가역적 물리 행동이 가능하며, VLA는 시각 장면·텍스트 지시 조작으로 위험을 더욱 악화시킬 수 있음
IDCardHuman audit
RAI4-0196
체화형 LLM 맥락적 탈옥
Embodied LLM contextual jailbreak
공격자가 피지컬 작업 맥락으로 프레임을 구성하여 embodied LLM 에이전트가 안전 제한을 우회하고 악의적 피지컬 행동을 수용하도록 유도하는 위험.
An attacker frames a physical task context so an embodied LLM agent bypasses safety restrictions and accepts malicious physical action instructions.
① Description
② L3 mapping
③ Duplicate
RAI4-0742
LLM 악의적 사용의 탈옥 - 백도어 공격
Jailbreak in LLM malicious use - backdoor attack
공격자가 학습 데이터셋에 백도어를 남겨 LLM이 평균적으로는 안전해 보이지만 특정 조건에서 유해한 콘텐츠를 생성하고, 이러한 백도어 행동이 여러 보안 학습 기법 적용 이후에도 지속되는 리스크
The risk that holes left in the training dataset make LLMs appear safe on average yet generate harmful content under specific conditions, a backdoor attack whose behaviours persist even after multiple security training techniques are applied.
Source members (2)
Source: min_cos=0.8492
RAI4-0742LLM 악의적 사용의 탈옥 - 백도어 공격
RAI4-0743LLM 악의적 사용의 탈옥 - 교육 데이터 오염
① Description
② L3 mapping
③ Duplicate
RAI4-0744
LLM 악의적 사용 시 탈옥 - 프롬프트 공격
Jailbreak in LLM malicious use - prompt attacks
프롬프트 및 추론 단계에서 프롬프트 주입, 역할극, 적대적 프롬프팅, 프롬프트 형식 변환 등 대화가 LLM을 혼란스럽거나 과도하게 순응하는 상태로 밀어 넣어 유해한 질문에 유해한 출력을 산출할 위험이 커지는 리스크
The risk that, in the prompting and reasoning phase, dialog such as prompt injection, role play, adversarial prompting, and prompt form transformation pushes LLMs into confused or overly compliant states, raising the risk of producing harmful outputs when confronted with harmful questions.
① Description
② L3 mapping
③ Duplicate
RAI4-0745
화이트박스·블랙박스 LLM 탈옥 공격
White-box and black-box LLM jailbreak attacks
미세조정 및 정렬 단계에서 정교하게 설계된 지시 데이터셋으로 LLM을 미세조정하여, 모델 가중치를 수정하는 화이트박스 공격과 API 기반 블랙박스 공격 모두에서 유해하거나 윤리 규범을 위반하는 콘텐츠를 생성하는 탈옥이 이루어지는 리스크
The risk that elaborately designed instruction datasets are used in the fine-tuning and alignment phase to drive LLMs to generate harmful content or content violating ethical norms, achieving a jailbreak through either white-box modification of parameter weights or black-box fine-tuning.
① Description
② L3 mapping
③ Duplicate
RAI4-0746
탈옥 공격
Jailbreak attacks
공격자가 모델에 설정된 가드레일을 뚫고 제한된 작업을 수행하게 만드는 리스크
The risk that a jailbreaking attack breaks through the guardrails established in the model to perform restricted actions.
Source members (2)
Source: min_cos=0.8613
RAI4-0746탈옥 공격
RAI4-1350탈옥 취약성
① Description
② L3 mapping
③ Duplicate
RAI4-0759
프롬프트 유출에 의한 기밀 지시문 노출
Exposure of confidential instructions via prompt leaking
공격자가 프롬프트 주입을 통해 모델이 사전 설계된 지시문을 출력하도록 오도하여, 비공개 프롬프트에 담긴 세부 정보와 LLM 애플리케이션의 핵심인 기밀 지시문이 노출되는 리스크
The risk that prompt leaking, a type of prompt injection attack, misleads a model into printing its pre-designed instructions, exposing details contained in private prompts and the confidential instructions central to LLM applications.
① Description
② L3 mapping
③ Duplicate
RAI4-0767
기술적 안전조치 우회 공격
Circumvention of technical safety measures
공격자가 모델의 취약점을 악용해 오용 완화를 위한 기술적 안전조치를 우회하고 모델과 그 역량에 무단 접근함으로써 의도치 않은 유해 행동을 유발하는 리스크
The risk that attackers exploit model vulnerabilities to circumvent the technical measures intended to mitigate misuse, gaining unauthorized access to a model and its capabilities and eliciting unwanted behaviour.
① Description
② L3 mapping
③ Duplicate
RAI4-0769
프롬프트 주입에 의한 시스템 장악
System compromise through prompt injection
공격자가 LLM 기반 대화형 시스템이나 그것이 검색할 데이터에 악의적 프롬프트를 삽입하여, 의도치 않은 동작 수행, 민감정보 공개, 원격 제어, 데이터 절취, 서비스 거부 등 시스템 장악을 초래하는 리스크
The risk that maliciously inserted prompts, injected directly or indirectly into data an LLM-based system retrieves, cause unintended actions, disclosure of sensitive information, remote control, data theft, or denial of service.
① Description
② L3 mapping
③ Duplicate
RAI4-0773
비텍스트 모달리티를 통한 LLM 공격
Attacks on LLMs through non-text modalities
이미지·영상 등 텍스트 외 모달리티를 처리하는 다중모달 LLM에서, 공격자가 이미지에 탈옥 텍스트나 미세한 교란을 삽입하여 안전장치를 우회하고 데이터 유출을 유발하는 리스크
The risk that attackers exploit multimodal LLMs by embedding jailbreaking text or imperceptible perturbations into images or video frames, bypassing safety mechanisms and enabling exfiltration attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-0782
딥러닝 프레임워크 취약점 악용
Exploitation of deep-learning framework vulnerabilities
LLM이 구현 기반으로 삼는 딥러닝 프레임워크의 버퍼 오버플로, 메모리 손상, 입력 검증 결함 등 취약점이 악용되어 모델과 시스템이 침해되는 리스크
The risk that vulnerabilities in the deep-learning frameworks underlying LLMs, such as buffer overflows, memory corruption, and input-validation issues, are exploited to compromise models and systems.
① Description
② L3 mapping
③ Duplicate
RAI4-0784
외부 도구 악용에 의한 정보 유출·주입 공격
Information leakage and injection via compromised external tools
적대적 도구 제공자가 API나 프롬프트에 악의적 지시를 삽입하여 LLM이 학습 데이터나 이용자 프롬프트의 민감정보를 유출하게 하고, 검증되지 않은 외부 입력이 주입 공격과 임의 코드 실행으로 이어지는 리스크
The risk that adversarial tool providers embed malicious instructions in APIs or prompts, causing LLMs to leak memorized sensitive information from training data or user prompts, while unverified external inputs enable injection attacks and arbitrary code execution.
① Description
② L3 mapping
③ Duplicate
RAI4-0786
추출 공격
Extraction attacks
공격자가 블랙박스 피해 모델을 질의하고 그 질의·응답으로 학습하여 거의 동일한 성능의 대체 모델을 구축하거나 LLM의 도메인 지식을 끌어낸 특화 모델을 개발하는 리스크
The risk that an adversary queries a black-box victim model and builds a substitute model by training on the queries and responses, achieving nearly the same performance or developing a domain-specific model that draws domain knowledge from the LLM.
① Description
② L3 mapping
③ Duplicate
RAI4-0789
GPU 부채널 공격에 의한 모델 파라미터 탈취
Model-parameter theft via GPU side-channel attacks
LLM 학습에 필요한 대규모 GPU 자원이 부채널 공격의 표면이 되어, 공격자가 학습된 모델의 파라미터를 추출하는 리스크
The risk that the significant GPU resources required to train LLMs introduce a side-channel attack surface through which adversaries extract the parameters of trained models.
① Description
② L3 mapping
③ Duplicate
RAI4-0917
개발 언어 인터프리터 취약점 위협
Programming language interpreter vulnerability threats
LLM 개발이 의존하는 Python 등 언어 인터프리터의 취약점이 개발된 모델의 보안을 위협하는 리스크.
The risk that vulnerabilities in language interpreters such as Python, on which LLM development depends, threaten the security of the developed models.
① Description
② L3 mapping
③ Duplicate
RAI4-0919
전처리 도구 취약점 악용
Exploitation of pre-processing tool vulnerabilities
LLM 파이프라인에서 사용되는 전처리 도구(예: OpenCV)의 취약점을 악용한 공격으로 시스템이 손상되는 리스크.
The risk that attacks exploiting vulnerabilities in pre-processing tools used in LLM pipelines, such as OpenCV, compromise the system.
① Description
② L3 mapping
③ Duplicate
RAI4-0921
메모리 취약점 기반 모델 파라미터 변조
Model parameter tampering via memory vulnerabilities
로우해머 등 메모리 관련 하드웨어 취약점이 악용되어 LLM의 파라미터가 변조되는(예: Deephammer 공격) 리스크.
The risk that memory-related hardware vulnerabilities such as rowhammer are leveraged to manipulate LLM parameters, as in Deephammer-style attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-0927
LLM 대상 신종 공격
Novel attack vectors against LLMs
프롬프트 추상화를 통한 API 비용 악용, RLHF 과정의 보상모델 백도어, LLM을 활용한 적대적 샘플 생성 등 신종 공격 기법이 LLM 시스템을 위협하는 리스크.
The risk that novel attack techniques—prompt abstraction attacks exploiting API pricing, backdoor attacks on the RLHF reward model, and LLM-based construction of adversarial samples—threaten LLM systems.
① Description
② L3 mapping
③ Duplicate
RAI4-0928
프롬프트 주입을 통한 목표 탈취
Goal hijacking via prompt injection
입력에 기존 지시를 무시하라는 문구를 주입하여 LLM에 설계된 원래 목표가 탈취되고 주입된 새 목표가 실행되는 리스크.
The risk that injecting a phrase such as ignore the above instruction and do into the input hijacks the original goal of the designed prompt in an LLM and executes the attacker's injected goal instead.
① Description
② L3 mapping
③ Duplicate
RAI4-0929
원스텝 탈옥
One-step jailbreaks
역할극 시나리오 설정, 양성 정보 통합, 난독화 등 프롬프트 자체를 직접 수정하는 단일 단계 기법으로 입출력 필터와 안전장치가 우회되는 리스크.
The risk that one-step jailbreaks—direct modifications to the prompt such as role-playing scenarios, benign-content integration, and obfuscation—circumvent input and output filters and safety measures.
① Description
② L3 mapping
③ Duplicate
RAI4-0930
다단계 탈옥
Multi-step jailbreaks
일련의 대화에서 요청을 단계적으로 맥락화하거나 외부 인터페이스·모델의 도움을 받아 시나리오를 구성함으로써 LLM이 유해하거나 민감한 콘텐츠를 단계적으로 생성하게 되는 리스크.
The risk that multi-step jailbreaks, constructing well-designed scenarios across a series of conversations through request contextualizing or external assistance, guide LLMs to generate harmful or sensitive content step by step.
① Description
② L3 mapping
③ Duplicate
RAI4-0943
보안 - 견고성
Security - robustness
프롬프트 인젝션·시각적 적대 예제 등 탈옥 기법과 백도어·모델 포이즈닝으로 안전 가드레일이 우회되고 모델·프롬프트가 탈취되는 등 시스템에 가해지는 위협이 실현되는 리스크.
The risk that threats posed to generative AI systems materialize through jailbreaking techniques such as prompt injection and visual adversarial examples, backdoors and model poisoning that bypass safety guardrails, and model or prompt theft.
① Description
② L3 mapping
③ Duplicate
RAI4-1058
악성코드 개발 비용 절감 지원
Lowering the cost of malware development
LM 기반 보조 코딩 도구가 탐지를 회피하도록 기능을 바꾸는 다형성 악성코드의 개발 비용을 낮추는 리스크.
The risk that assistive coding tools based on LMs lower the cost of developing polymorphic malware able to change its features in order to evade detection.
① Description
② L3 mapping
③ Duplicate
RAI4-1173
역할극 지시 악용
Role-play instruction exploitation
공격자가 입력 프롬프트에서 모델에 과격주의자나 인종차별주의자처럼 위험한 집단과 연관된 역할 속성을 부여하고 지시를 내려, 지시에 지나치게 충실한 모델이 그 인물의 말투로 안전하지 않은 콘텐츠를 산출하는 리스크.
The risk that attackers specify a model's role attribute within the input prompt, tying it to potentially risky groups such as radicals or racial discriminators, so that the overly faithful model outputs unsafe content in the style of that character.
① Description
② L3 mapping
③ Duplicate
RAI4-1176
역노출
Reverse exposure
공격자가 모델이 금지된 출력을 생성하도록 유도하여 불법·비윤리 정보에 접근함으로써 안전 통제가 접근 통로로 역전되는 리스크
Attackers induce the model to generate prohibited outputs and thereby extract illegal or unethical information, reversing safety controls into an access channel.
① Description
② L3 mapping
③ Duplicate
RAI4-1314
공격적 사이버 역량 악용
Offensive cyber capability misuse
LLM이 하드웨어와 소프트웨어, 데이터의 취약점을 탐지하고 악용하며 시스템이나 네트워크 내부에서 탐지를 회피한 채 특정 목표 달성에 집중하는 사이버 역량을 갖추어 공격에 쓰이는 리스크.
The risk that an LLM possesses cyber-domain capabilities to detect and exploit vulnerabilities in hardware, software, and data and to evade detection once inside a system or network while focusing on achieving specific objectives.
① Description
② L3 mapping
③ Duplicate
RAI4-1348
대규모 표적 괴롭힘
Targeted harassment at scale
LLM이 온라인에서 개인을 표적으로 배포되어 개인화된 유해 메시지를 대규모로 발송함으로써 표적 괴롭힘이 발생하는 리스크.
The risk that LLMs are deployed to target individuals online, sending them personalized and harmful messages at scale.
① Description
② L3 mapping
③ Duplicate
RAI4-1354
대규모 사이버범죄 오용
Cybercrime misuse at scale
악의적 행위자가 생성 모델을 탈옥시켜 민감·유해 콘텐츠와 표적 맞춤형 설득 자료를 생산함으로써 사이버범죄를 저비용·대규모로 수행하는 리스크
Malicious actors jailbreak generative models to produce sensitive or harmful content and generate persuasive, individually tailored material, conducting cybercrime efficiently at scale and reduced cost.
① Description
② L3 mapping
③ Duplicate
RAI4-1513
해석 가능성 기술의 오용
Misuse of interpretability techniques
모델에 대한 이해를 높이는 해석가능성 기법이 안전 관련 특징을 부호화한 뉴런의 활성 저하나 정보 검열에 사용되거나 화이트박스 공격 시나리오 모의와 적대적 공격 개발에 활용되는 리스크.
The risk that interpretability techniques, by enabling a better understanding of the model, are used for harmful purposes such as identifying and modifying neurons that encode safety-related features to decrease their activation or censor information, or simulating a white-box attack scenario to aid the development of adversarial attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-1517
의도된 동작을 전복하는 모델 탈옥
Jailbreak of a model to subvert intended behavior
내부 파라미터 접근이 필요한 화이트박스 자동 생성, 모델 내부에 접근하지 않는 블랙박스 방식, 추론·역할극을 이용한 사람이 읽을 수 있는 프롬프트 등으로 배포 중 모델에 적대적 입력이 주입되어 안전 메커니즘이 우회되고 의도된 사용에서 벗어난 모델 동작이 발생하는 리스크.
The risk that adversarial inputs supplied to a model during deployment result in model behavior deviating from intended use. Such jailbreaks may be generated automatically in white-box settings requiring access to internal training parameters, crafted in black-box settings without access to model internals, or written as human-readable prompts using reasoning or role-play to convince the model to bypass its safety mechanisms.
① Description
② L3 mapping
③ Duplicate
RAI4-1518
멀티모달 모델 탈옥(jailbreak)
Jailbreak of a multimodal model
멀티모달 GPAI 모델에 대한 적대적 탈옥 공격이 높은 성공률로 임의의 또는 특정 출력을 유도하고 컨텍스트 창 등 모델 내부 정보를 유출시키는 리스크.
The risk that adversarial jailbreak attacks on current-generation multimodal GPAI models automatically induce arbitrary or specific outputs with high success rates and exfiltrate the model's context window or other model internals.
① Description
② L3 mapping
③ Duplicate
RAI4-1520
저빈도 인코딩·저자원 언어를 통한 안전 훈련 우회
Safety-training bypass via rare encodings and low-resource languages
Base64 등 안전 미세조정에 충분히 포함되지 않은 텍스트 인코딩이나 저자원 언어로 유해 프롬프트를 변환하여 모델의 안전장치를 우회하는 리스크.
The risk that harmful natural language prompts translated into text encodings such as Base64 or into low-resource languages that safety fine-tuning covers little or not at all are used to craft jailbreak attacks that bypass a model's safeguards.
① Description
② L3 mapping
③ Duplicate
RAI4-1521
추가 모달리티로 인한 공격면 확대
Attack-surface expansion from additional modalities
멀티모달 모델에서 모달리티별 견고성 차이로 공격자가 가장 취약한 모달리티를 선택할 수 있어 새로운 공격 벡터가 생기고 탈옥부터 데이터 오염까지 기존 공격의 범위가 확대되는 리스크.
The risk that additional modalities introduce new attack vectors in multimodal models and expand the scope of previous attacks ranging from jailbreaking to poisoning, as differing robustness levels across modalities let malicious actors choose the most vulnerable part of the model to attack.
① Description
② L3 mapping
③ Duplicate
RAI4-1522
다수 예시 장문맥 탈옥
Many-shot long-context jailbreaking
긴 컨텍스트 창에 다수의 유해 출력 예시를 제시하여 짧은 컨텍스트 모델에서는 통하지 않던 공격이 성공하고 유해 응답이 유도되며, 컨텍스트 창이 확대될수록 이러한 취약성이 커지는 리스크.
The risk that language models with long context windows are exploited by many-shot jailbreaking, where supplying a high number of examples of the desired harmful output elicits undesirable responses that few-shot attacks fail to trigger, with the vulnerability becoming more significant as context windows expand.
① Description
② L3 mapping
③ Duplicate
RAI4-1652
LLM 이중용도 역량의 악용·오용
Malicious use and misuse of dual-use LLM capabilities
악의적 행위자가 LLM의 이중용도 역량을 악용·오용하여 피해가 발생하는 리스크.
The risk that malicious actors misuse the dual-use capabilities of LLMs to cause harm.
① Description
② L3 mapping
③ Duplicate
RAI4-1654
LLM 맞춤형 피싱에 의한 사회공학 공격 증폭
Amplified social engineering through LLM-crafted tailored phishing
LLM이 대규모로 개인 맞춤형 피싱 메시지를 작성하여 이용자가 민감 정보를 제공하거나 공격자에게 접근 권한을 내주도록 유도되고, 이러한 사회공학 공격이 대규모 해킹 작전의 기반이 되는 리스크.
The risk that LLMs craft personalized phishing emails or messages at scale that are harder for users to recognize, tricking them into disclosing sensitive information or granting adversary access to critical resources and forming the basis of larger hacking operations.
① Description
② L3 mapping
③ Duplicate
RAI4-1664
탈옥·프롬프트 주입에 의한 LLM 보안 실패
Security failures from jailbreaks and prompt injection in LLMs
LLM이 적대적으로 견고하지 않고 입력 내 권한 수준 구분이 없어 탈옥과 프롬프트 주입 공격에 취약하며, 표준화된 평가와 효율적 화이트박스 검증 방법의 부재로 이러한 보안 실패를 제거하기 어려운 리스크.
The risk that LLMs, lacking adversarial robustness and robust privilege levels within their input, remain vulnerable to jailbreak and prompt-injection attacks that are hard to eliminate given the absence of standardized evaluation and efficient white-box robustness assessment.
① Description
② L3 mapping
③ Duplicate
RAI4-1665
페르소나 지정·사회공학 기법에 의한 안전장치 우회
Safeguard bypass through persona assignment and social-engineering tricks
공격자가 모델에 특정 페르소나를 지정하거나 인간 또는 다른 LLM이 고안한 사회공학적 기법 등 심리적 속임수를 사용하여 모델을 악용하는 리스크.
The risk that attackers exploit psychological tricks on LLMs, such as instructing the model to behave like a specific persona or employing social-engineering techniques crafted by humans or other LLMs.
① Description
② L3 mapping
③ Duplicate
RAI4-1666
프록시 목표 최적화를 통한 탈옥 자동 탐색
Automated discovery of jailbreaks through proxy-objective optimization
공격자가 탈옥 성공과 잡음 있게 상관된 프록시 목표에 대해 수동 또는 자동으로 그래디언트 기반 및 비그래디언트 적대적 최적화를 수행하여 탈옥 공격을 발견하는 리스크.
The risk that jailbreak attacks are discovered by performing manual or automated, gradient-based or gradient-free adversarial optimization against a proxy objective noisily correlated with jailbreak success.
① Description
② L3 mapping
③ Duplicate
RAI4-1676
신체적 피해로 이어지는 체화형 탈옥
Embodied jailbreak to physical harm
체화된 LLM에 대한 탈옥이 손상된 추론을 되돌릴 수 없는 실제 물리적 행동으로 전환시켜 상해나 손해가 발생하는 리스크.
The risk that a jailbreak of an embodied LLM translates compromised reasoning into irreversible real-world physical actions, such as manipulation causing injury or damage.
① Description
② L3 mapping
③ Duplicate
RAI4-1686
궤적 지속형 적대적 공격
Trajectory-persistent adversarial attack
월드 모델 인코더에 가해진 단일 적대적 교란이 순환 잠재 상태를 통해 전파되어 다단계 롤아웃 전체를 손상시키며, 무상태 모델보다 훨씬 파괴적으로 초기 단계에서 증폭되는 리스크.
The risk that a single adversarial perturbation to a world-model encoder propagates through recurrent latent state, corrupting an entire multi-step rollout far more destructively than in a stateless model through early-step amplification.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-02 물리적 공격 Physical Attacks18 cards
직접적인 하드웨어 변조를 통해 구성 요소를 조작하거나 성능을 방해하고 물리적 손상을 가하는 공격으로, 로봇의 안전 기능이 무력화될 수 있음
IDCardHuman audit
RAI4-0197
유해 행동 재구성 공격
Harmful-action reframing attack
공격자가 유해한 물리 행동 지시를 무해하거나 정상적인 작업 요청처럼 바꾸어 embodied 에이전트가 이를 수용·실행하도록 유도하는 위험.
An attacker rephrases a harmful physical instruction as a benign or task-compliant request so that the embodied agent accepts and executes it.
① Description
② L3 mapping
③ Duplicate
RAI4-0198
완곡 표현을 이용한 유해 의도 은폐
Euphemistic concealment of harmful physical intent
공격자가 유해한 대상과 결과는 유지한 채 완곡어·대체 표현·간접 개념으로 물리적 위해 의도를 숨기는 위험.
An attacker conceals harmful physical intent with euphemisms, substitutions, or indirect concepts while preserving the same harmful target and outcome.
① Description
② L3 mapping
③ Duplicate
RAI4-0199
명시적 악의적 물리 요청 실행
Execution of explicitly malicious physical requests
embodied 에이전트가 신체 위해·침입·절도·파괴 등 불법적인 물리 행동을 명시적으로 요구한 사용자 요청을 수용하고 실행하는 위험.
An embodied agent accepts and begins executing a user request that explicitly calls for physical harm, intrusion, theft, sabotage, or another illegal physical act.
① Description
② L3 mapping
③ Duplicate
RAI4-0201
피지컬 AI 파괴 행위
Embodied sabotage
피지컬 AI 시스템이 장비·인프라 또는 다른 피지컬 시스템을 손상·무력화·방해·변조하도록 유도되는 위험.
A physical AI system is induced to damage, disable, obstruct, or tamper with equipment, infrastructure, or other physical assets.
① Description
② L3 mapping
③ Duplicate
RAI4-0203
피지컬 AI의 혐오·학대 행동
Hateful or abusive embodied action
피지컬 AI 시스템이 사람들을 향한 차별적·괴롭힘·위협·학대적 행동을 수행하도록 유도되는 위험.
A physical AI system is directed to perform discriminatory, harassing, intimidating, or abusive actions toward people in shared environments.
① Description
② L3 mapping
③ Duplicate
RAI4-0210
인지-행동 순서 조작 공격
Perception-action sequence manipulation attack
공격자가 영상 순서의 관측이나 행동 단서를 조작해 로봇이 명시된 충돌·힘·이격거리·대상물 사용 제약을 위반하게 하는 위험.
An attacker alters observations or action cues across a video sequence so that a robot violates a specified collision, force, distance, or object-use constraint.
① Description
② L3 mapping
③ Duplicate
RAI4-0221
로봇 인지 겨냥 적대 패치 공격
Adversarial patch attack on robot perception
공격자가 물체나 표지에 최적화된 시각 패턴을 부착해 특정 인지 오류와 후속 위험 행동을 유발하는 리스크.
The risk that an attacker places an optimized visual pattern on an object or sign to induce a targeted perception error and downstream unsafe robot action.
① Description
② L3 mapping
③ Duplicate
RAI4-0223
언어 명령 기반 로봇 정책 공격
Language-instruction attack on robot policy
악의적 언어 지시·접미사·프롬프트가 의도된 로봇 정책을 변경하여 비안전 피지컬 행동을 유발하는 위험.
Malicious language instructions, suffixes, or prompts alter the intended robot policy and induce unsafe physical behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0302
조작·이동 능력 공격 전용화
Manipulation/mobility repurposed for attack
조작 및 이동 기능이 타격·투척·돌진 공격에 전용되는 위험.
A robot’s manipulation or mobility capabilities are repurposed to physically attack people, property, or infrastructure.
① Description
② L3 mapping
③ Duplicate
RAI4-0305
자동화된 표적 스토킹·물리적 위협
Automated targeted stalking and physical intimidation
운영자나 공격자가 특정인을 위협하거나 괴롭힐 목적으로 로봇에 반복 추적·진로 차단·고립·접근을 지시하는 위험.
An operator or attacker directs a robot to repeatedly follow, block, corner, or approach a specific person in order to intimidate or harass them.
① Description
② L3 mapping
③ Duplicate
RAI4-0311
피지컬 행동 편향 실행
Bias executed as physical behavior
기반 모델 편향이 분류·회피·차별적 서비스 등 피지컬 행동으로 실행되는 위험.
A robot turns biased model outputs into unequal physical service, movement, or treatment.
① Description
② L3 mapping
③ Duplicate
RAI4-0323
장면 의미 조작 공격
Semantic scene manipulation attack
공격자가 모델이나 센서를 직접 변경하지 않고 표지·물체 배치·의복 패턴·장면 맥락을 바꾸어 로봇이 환경을 잘못 해석하게 만드는 리스크.
The risk that an attacker changes a sign, object placement, clothing pattern, or scene context so that a robot misinterprets the environment without modifying the model or sensor.
① Description
② L3 mapping
③ Duplicate
RAI4-0347
프롬프트→행동 주입 공격
Prompt-to-act injection
언어 매개 로봇 또는 에이전트가 악의적 지시·표지·QR코드·음성 명령·문서를 안전 위반 피지컬 행동으로 변환하는 위험.
A language-mediated robot or agent may convert malicious instructions, signs, QR codes, voice commands, or documents into unsafe physical actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0756
프롬프트 역산 공격
Prompt inversion attack
공격자가 AI의 출력이나 관찰 가능한 행태로부터 비공개 또는 독점 프롬프트 문구를 권한 없이 복원하는 리스크
The risk that a prompt-inversion attack recovers private or proprietary prompt text from an AI system's outputs or observable behavior without authorization.
① Description
② L3 mapping
③ Duplicate
RAI4-0758
프롬프트 인젝션 공격
Prompt injection attack
적대적 입력의 한 형태로서 공격자가 생성형 AI 시스템에 주어지는 텍스트 지시를 조작하고 시스템 지시와 이용자 데이터가 분리되지 않은 아키텍처의 허점을 악용하여 유해한 출력을 산출하게 하며, 서비스 거부나 AI 탐지 우회를 일으키는 리스크.
The risk that prompt injections, a form of adversarial input, manipulate the text instructions given to a GenAI system and exploit architectural loopholes with no separation between system instructions and user data to produce harmful output, including flooding a model to cause denial-of-service attacks or to bypass AI detection software.
① Description
② L3 mapping
③ Duplicate
RAI4-1199
프롬프트 공격
Prompt attacks
정교하게 통제된 적대적 섭동이 텍스트 분류에서 모델의 답변을 뒤집고, 질문을 특정 방식으로 비틀어 모델이 답하지 않기로 한 위험 정보를 이끌어내는 리스크.
The risk that carefully controlled adversarial perturbation flips a model's answer when used to classify text inputs, and that twisting the prompting question solicits dangerous information the model chose not to answer.
① Description
② L3 mapping
③ Duplicate
RAI4-1675
언어 거부에도 실행되는 유해 물리 행동
Harmful action executed despite verbal refusal under action-space misalignment
체화된 LLM이 언어 출력 공간과 행동 출력 공간 간 정렬 불량으로 인해 유해한 요청을 언어로는 거부하면서 해당 물리적 행동은 그대로 실행하는 리스크.
The risk that an embodied LLM verbally refuses a harmful request while still executing the corresponding physical action, due to misalignment between its linguistic and action output spaces.
① Description
② L3 mapping
③ Duplicate
RAI4-1677
개념적 속임수에 의한 유해 물리 행동 유도
Harmful physical action induced by conceptual deception
유해한 물리적 과업을 무해한 개념적 용어로 재구성함으로써 체화 에이전트가 유해성을 인식하지 못한 채 해당 행동을 수행하는 리스크.
The risk that reframing a harmful physical task in benign conceptual terms induces an embodied agent to perform unrecognized harmful behavior.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-03 사이버보안 위협 Cybersecurity Threats29 cards
IoT·클라우드 인프라와의 통합 증가로 인해 광범위한 사이버 공격에 노출됨. Mirai 봇넷처럼 취약한 인증을 악용한 DDoS 공격, 드론 GPS 스푸핑을 통한 경로 하이재킹 등이 대표적 위협임
IDCardHuman audit
RAI4-0441
자동화된 사기 콘텐츠
Automated fraud content
AI 시스템이 피싱, 사기, 가짜 리뷰, 공문서 위장 자료의 생산을 대규모화하여 자동화 사기의 비용을 낮추고 도달 범위를 확대하는 리스크
AI systems scale the production of phishing messages, scams, fake reviews, and fraudulent official-looking material, lowering the cost and increasing the reach of automated fraud.
① Description
② L3 mapping
③ Duplicate
RAI4-0448
맞춤형 AI 사기
Personalized AI scams
AI가 표적별로 개인화된 설득력 있는 사기 콘텐츠를 대규모로 생성하여 금융·사회공학 사기의 성공률을 높이는 리스크
AI generates persuasive scam content individually tailored to each target at scale, raising success rates of financial and social-engineering fraud.
① Description
② L3 mapping
③ Duplicate
RAI4-0620
AI 기반 사이버 공격 확대
AI-enabled cyberattack escalation
AI가 취약점 발견·악용, 비밀번호 크래킹, 악성코드 생성, 정교한 피싱, 네트워크 스캐닝, 사회공학을 자동화·고도화하여 공격 진입 장벽을 낮추고 방어 복잡도를 높여 핵심 인프라 마비와 대규모 정보 유출 및 상당한 경제적 손실을 초래하는 리스크
The risk that AI automates and enhances vulnerability discovery and exploitation, password cracking, malicious code generation, sophisticated phishing, network scanning, and social engineering, lowering the barrier to entry for attackers while increasing the complexity of defense and leading to critical infrastructure paralysis, widespread data breaches, and substantial economic losses.
Source members (4)
Source: min_cos=0.8175 · Mixed L3
RAI4-0445사이버 공격 자동화
RAI4-0609AI 기반 공격적 사이버 작전
RAI4-0620AI 기반 사이버 공격 확대
RAI4-1340사이버 공격 목적 AI 남용
① Description
② L3 mapping
③ Duplicate
RAI4-0621
AI에 의한 사이버 공격 역량 증강
AI-amplified cyber offense capability
기존 사이버 위협이 AI로 인해 악화되어 LLM 에이전트 팀이 제로데이 취약점 악용까지 수행하게 되고 사이버전이 치명적 피해의 신뢰할 만한 위협이 되는 리스크
The risk that AI exacerbates existing cyber threats, with teams of LLM agents able to exploit zero-day vulnerabilities, making cyberwarfare a credible threat of catastrophic harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0648
AI 조력 취약점 탐색의 공격 문턱 저하
Lowered attack barriers from AI-assisted vulnerability discovery
AI 보조 도구가 소프트웨어 취약점 식별과 익스플로잇 코드 작성을 자동화하여 전문 지식이 필요하던 공격적 사이버 작전의 진입 문턱을 낮추고 초심자까지 무단 접근과 통제 획득을 가능하게 하는 리스크
The risk that AI assistants automate the identification of software vulnerabilities and the creation of exploit code, lowering the barrier to offensive cyber operations that previously required specialist programming knowledge and enabling novices to gain unauthorized access or control.
① Description
② L3 mapping
③ Duplicate
RAI4-0649
AI 기반 대규모 스피어피싱
AI-powered spear-phishing at scale
공격자가 AI로 신뢰할 수 있는 주체의 소통 양식을 모방한 고도로 개인화된 피싱 메시지를 대규모로 생성하고 긴급성과 공포를 자극하여, 피해자로부터 민감 정보를 탈취하거나 유해한 행동을 유도하는 리스크
The risk that attackers leverage AI to craft highly convincing, personalized spear-phishing messages at scale that imitate trusted entities and exploit urgency and fear, extracting sensitive information from victims or luring them into harmful actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0654
범용 AI에 의한 사이버 공격 증폭
AI amplification of cyberattack scale and effectiveness
범용 AI 모델이 악의적 행위자의 기존 역량과 자원을 증폭시켜 취약점 자동 탐색, 익스플로잇의 유연한 대규모 적용, 공격 기획·정찰·원격 통제·악성코드 구현·데이터 유출 지원과 사회공학 결합을 가능하게 하여 사이버 공격의 규모와 효과가 크게 확대되는 리스크
The risk that general-purpose AI models significantly enhance the magnitude and effectiveness of cyberattacks by amplifying malicious actors' existing capabilities and resources, automating vulnerability scanning, applying known exploits flexibly at scale, assisting with planning, reconnaissance, remote control, malware implementation, and data exfiltration, and combining social engineering with attacks at scale.
① Description
② L3 mapping
③ Duplicate
RAI4-0669
AI 조력 악성코드 생성
AI-assisted malicious code generation
LLM이 매우 낮은 비용과 빠른 속도로 양질의 코드를 작성하는 능력이 악용되어 공격자가 악성 공격 코드를 작성하고 사이버 공격을 자동화하는 리스크
The risk that the ability of LLMs to write reasonably good-quality code at extremely low cost and incredible speed is leveraged by malicious hackers to assist with cyberattacks and automate them.
Source members (4)
Source: min_cos=0.6905 · Mixed L3
RAI4-0668LLM 조력 사이버 공격 자동화
RAI4-0669AI 조력 악성코드 생성
RAI4-1135악성코드 생성
RAI4-1203LLM 조력 악성코드 개발
① Description
② L3 mapping
③ Duplicate
RAI4-0934
악의적 행위자의 AI 오용 조력
AI empowerment of malicious actors
음성 복제·딥페이크 생성, 패스워드 크래킹 가속, 비숙련자의 익스플로잇·피싱 제작 지원 등 AI가 악의적 행위자의 유해 행위를 가능하게 하거나 증폭하는 리스크.
The risk that AI enables or amplifies malicious actors' harmful actions, including voice cloning and deepfake creation, accelerated password cracking, and enabling unskilled actors to produce software exploits and effective phishing emails.
① Description
② L3 mapping
③ Duplicate
RAI4-0946
AI 기반 사이버 범죄
AI-enabled cybercrime
생성형 AI가 인간 사칭, 가짜 신원 생성, 음성 복제, 피싱 메시지 제작 등 사회공학 공격과 악성코드 생성·해킹에 오용되는 리스크.
The risk that generative AI is misused for fraudulent online activities, including social engineering attacks via human impersonation, fake identities, voice cloning, and phishing, and for generating malicious code or hacking.
Source members (2)
Source: min_cos=0.8344 · Mixed L3
RAI4-0946AI 기반 사이버 범죄
RAI4-1210생성형 AI 데이터 보안 노출
① Description
② L3 mapping
③ Duplicate
RAI4-0962
AI 조력 디지털 범죄
AI-supported digital crime
AI가 악성코드 제작과 해킹 등 디지털 범죄의 수행을 지원·가능하게 하는 리스크.
The risk that AI supports and enables the perpetration of digital crimes, including AI-supported malware and hacking.
① Description
② L3 mapping
③ Duplicate
RAI4-0965
위협 행위자 역량의 가속적 증대
Accelerated threat-actor capability uplift
AI가 사이버 위협 행위자의 익스플로잇 실행 속도·파급력을 높이고 딥페이크 등 허위정보의 생성 속도와 효과를 가속하는 리스크.
The risk that AI enables cyber threat actors to execute exploits with greater speed and impact and to generate disinformation such as deepfake media at accelerated rates and effectiveness.
① Description
② L3 mapping
③ Duplicate
RAI4-1086
AI 기반 사기
AI-enabled fraud
AI가 사기, 부정행위, 위조, 사칭 사기를 촉진하는 리스크.
The risk that AI facilitates fraud, cheating, forgery, and impersonation scams.
① Description
② L3 mapping
③ Duplicate
RAI4-1134
공격적 사이버 작전 악용
Malicious use for offensive cyber operations
고급 AI 비서가 공격자에게 이용되어 공격을 자동화하고 시스템과 네트워크의 취약점을 식별·악용하며 피싱과 악성 페이로드, 악성코드를 생성하는 리스크.
The risk that advanced AI assistants are used by attackers in offensive cyber operations to automate attacks, identify and exploit weaknesses in systems and networks, and generate phishing and malicious code payloads.
① Description
② L3 mapping
③ Duplicate
RAI4-1138
사기성 서비스 대규모 생성
Fraudulent services at scale
마크업 생성과 외부 도구 통합이 가능한 AI 비서가 악의적 행위자의 사기성 웹사이트와 애플리케이션을 대규모로 만들어, 신용카드 번호와 계정 정보 같은 민감정보를 탈취하거나 추가 악성코드를 설치하는 리스크.
The risk that AI assistants able to produce markup and use external tools help malicious actors create fraudulent websites and applications at scale that harvest sensitive information such as credit card numbers and credentials or install additional malware.
① Description
② L3 mapping
③ Duplicate
RAI4-1242
실세계 자원 접근 확대와 자기증식
Expanded resource access and self-proliferation
미래 AI 시스템이 웹사이트와 실제 행동에 접근해 허위 정보를 퍼뜨리고 사용자를 기만하며 네트워크 보안을 교란하고 악의적 행위자에게 탈취되며, 데이터와 자원 접근 확대로 자기증식하여 실존적 위험을 낳는 리스크.
The risk that future AI systems gaining access to websites and real-world actions disseminate false information, deceive users, disrupt network security, are compromised by malicious actors, and use increased access to data and resources for self-proliferation, posing existential risks.
① Description
② L3 mapping
③ Duplicate
RAI4-1338
AI 컴퓨팅 인프라 보안 위협
AI computing infrastructure security threats
AI 훈련·운영을 뒷받침하는 컴퓨팅 인프라가 다양하고 유비쿼터스적인 컴퓨팅 노드와 자원에 의존함으로써 컴퓨팅 자원의 악의적 소비와 컴퓨팅 인프라 계층에서의 보안 위협 국경 간 전파에 노출되는 리스크.
The risk that the computing infrastructure underpinning AI training and operations, relying on diverse and ubiquitous computing nodes and various computing resources, is exposed to malicious consumption of computing resources and cross-boundary transmission of security threats at the computing infrastructure layer.
① Description
② L3 mapping
③ Duplicate
RAI4-1342
범죄 활동 조력 AI 오용
AI misuse assisting criminal activity
AI가 범죄 기술 교습, 불법 행위 은폐, 불법·범죄용 도구 제작 등을 통해 테러·폭력·도박·마약과 관련된 전통적 불법 및 범죄 활동에 사용되는 리스크.
The risk that AI is used in traditional illegal or criminal activities related to terrorism, violence, gambling, and drugs, such as teaching criminal techniques, concealing illicit acts, and creating tools for illegal and criminal activities.
① Description
② L3 mapping
③ Duplicate
RAI4-1347
신원 도용 목적 AI 사칭
AI-generated impersonation for identity theft
AI가 생성한 사칭이 신원 도용에 사용되어 개인에 대한 피해와 기만이 발생하는 리스크.
The risk that AI-generated impersonation is used for identity theft, producing deception and harm to the person.
Source members (3)
Source: min_cos=0.7809 · Mixed L3
RAI4-0449AI 신원 스푸핑
RAI4-1346합성 신원 악용
RAI4-1347신원 도용 목적 AI 사칭
① Description
② L3 mapping
③ Duplicate
RAI4-1355
사이버 공격 오용
Cyberattack misuse
생성형 AI가 표적 시스템의 핵심 취약점 식별과 새로운 침투 방법 발견을 통해 사이버 공격의 접근성·성공률·규모·속도·은밀성·파괴력을 높여, 전력망·금융 시스템·무기 관리 시스템 등 핵심 인프라에 심각한 피해가 발생하는 리스크.
The risk that generative AI amplifies the frequency and destructiveness of cyberattacks by increasing their accessibility, success rate, scale, speed, stealth, and potency, identifying critical vulnerabilities in targeted systems and discovering innovative methods of infiltration, inflicting significant damage on critical infrastructure including electrical grids, financial systems, and weapons management systems.
① Description
② L3 mapping
③ Duplicate
RAI4-1377
사이버 공격 장벽 저하·공격표면 확대
Lowered cyberattack barriers and expanded attack surface
취약점의 자동 발견과 악용 등으로 해킹·맬웨어·피싱 등 공격적 사이버 역량의 장벽이 낮아지고 표적 공격의 공격표면이 확대되어 시스템 가용성과 학습 데이터·코드·모델 가중치의 기밀성·무결성이 훼손되는 리스크.
The risk that lowered barriers for offensive cyber capabilities, including automated discovery and exploitation of vulnerabilities easing hacking, malware, phishing, and other cyberattacks, together with an increased attack surface for targeted cyberattacks, compromise a system's availability or the confidentiality or integrity of training data, code, or model weights.
① Description
② L3 mapping
③ Duplicate
RAI4-1391
범용 AI에 의한 사이버범죄 효율 상승
Cybercrime efficiency uplift by general-purpose AI
범용 AI 역량이 IT 기반 사기를 중심으로 사이버범죄의 효율과 효과를 높여 공격 규모와 성공률을 동시에 확대하는 리스크
General-purpose AI capabilities improve the efficiency and efficacy of cybercrime, especially IT-leveraged fraud, expanding both the scale and success rate of attacks.
① Description
② L3 mapping
③ Duplicate
RAI4-1399
자율 복제
Autonomous replication
AI 소프트웨어가 웜·바이러스와 유사하게 대응 조치에도 불구하고 네트워크를 통해 자율적으로 복제·확산되는 리스크
AI software autonomously replicates and spreads across networks despite countermeasures, in the manner of self-propagating worms and viruses.
① Description
② L3 mapping
③ Duplicate
RAI4-1408
AI 주도 취약점 발견·악용
AI-driven vulnerability discovery and exploitation
AI 시스템이 소프트웨어와 사이버 인프라의 취약점을 발견·악용하여 모델 역량이 공격적 보안 위협으로 전환되는 리스크
AI systems discover and exploit vulnerabilities in software and cyberinfrastructure, converting model capability into offensive security risk.
① Description
② L3 mapping
③ Duplicate
RAI4-1561
취약점 자동 탐지·악용에 의한 사이버공격 확대
Cyberattack scaling through automated vulnerability discovery and exploitation
GPAI가 소프트웨어 취약점의 자동 발견과 악성코드 자동 개발을 지원하여 악의적 행위자의 사이버공격이 저비용으로 대규모화되고 피해가 증대되는 리스크.
The risk that GPAI aids automated discovery of software vulnerabilities and automated malware development, allowing malicious actors to scale cyberattacks at low cost and increase their impact.
① Description
② L3 mapping
③ Duplicate
RAI4-1585
유해 작업 자동화에 의한 피해 규모 증폭
Amplified harm scale from automated harmful workflows
AI가 유해한 작업 흐름을 자동화하거나 확장하여 그 속도·도달 범위·지속성·표적 수가 크게 증가하는 리스크.
The risk that AI automates or expands a harmful workflow so that its speed, reach, persistence, or target count substantially increases.
① Description
② L3 mapping
③ Duplicate
RAI4-1612
프런티어 AI 오용에 의한 위협 행위자 역량 상승
Threat-actor capability uplift from frontier AI misuse
프런티어 AI가 사이버공격 수행, 허위정보 캠페인 운영, 생물·화학 무기 설계를 지원하여 정교하지 않은 위협 행위자의 진입 장벽이 낮아지는 리스크.
The risk that frontier AI helps bad actors perform cyberattacks, run disinformation campaigns, and design biological or chemical weapons, continuing to lower barriers to entry for less sophisticated threat actors.
① Description
② L3 mapping
③ Duplicate
RAI4-1614
AI 기반 사이버 작전
AI-enabled cyber operations
AI 프로그래밍 역량이 맞춤형 피싱과 악성코드 복제를 통해 사이버 작전을 더 빠르고 효과적이며 대규모로 만들고, 공격·방어 양측에서 인간 감독이 축소되는 리스크
AI programming capability scales cyber operations, enabling faster, more effective, and larger intrusions through tailored phishing and replicated malware, with declining human oversight on both offense and defense.
① Description
② L3 mapping
③ Duplicate
RAI4-1618
자율 사이버 공격을 통한 인간 통제 감소
Human control reduction through autonomous cyber offence
AI 시스템이 컴퓨터 시스템 취약점을 악용해 자금, 컴퓨팅, 핵심 인프라에 접근하고 궁극적으로 자율적 사이버 공격을 수행하여 인간 통제를 약화시키는 리스크
AI systems acquire influence by exploiting computer-system vulnerabilities, gaining access to money, compute, and critical infrastructure, and eventually executing cyberattacks autonomously to reduce human control.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-04 센서·입력 검증 실패 Sensor & Input Validation Failures22 cards
센서 오작동이나 입력 검증 부실로 인해 환경을 잘못 평가하고, 그 결과 안전하지 않은 물리 행동이 유발될 수 있음
IDCardHuman audit
RAI4-0044
맥락적 프라이버시 보호 실패
Contextual privacy protection failure
적절한 행동이 사회적·상황적 단서에 좌우되는 프라이버시 민감 상황에서 에이전트가 맥락적 개인정보를 보호하지 못하는 리스크.
The risk that an agent fails to protect contextual personal information in privacy-sensitive situations where the appropriate action depends on social and situational cues.
① Description
② L3 mapping
③ Duplicate
RAI4-0180
희귀 가정 상해의 전조 감지 실패
Failure to detect rare household injury precursors
가정용 에이전트가 드물지만 가능한 낙상·중독·화상·열상·압착 사고의 전조를 감지하지 못해 경고하거나 개입하지 않는 위험.
A household agent fails to detect early signs of rare but plausible falls, poisoning, burns, lacerations, or crush events and therefore does not warn or intervene.
① Description
② L3 mapping
③ Duplicate
RAI4-0181
현재 물리적 위험 상태 감지 실패
Failure to detect a present physical hazard
멀티모달 또는 embodied 시스템이 행동을 선택하기 전에 현재 텍스트·이미지·영상·센서 입력에 나타난 위험 상태를 식별하지 못하는 위험.
A multimodal or embodied system fails to identify a hazardous condition already visible in current text, image, video, or sensor input before selecting an action.
① Description
② L3 mapping
③ Duplicate
RAI4-0182
물리적 위험 개입 실패
Physical danger intervention failure
시스템이 위험한 물리적 상황을 인식하고도 적시에 적절한 개입이나 거부 응답을 생성하지 못하는 리스크.
The risk that a system recognizes a hazardous physical situation but fails to produce a timely and appropriate intervention or refusal.
① Description
② L3 mapping
③ Duplicate
RAI4-0224
제약 모니터링 실패
Constraint monitoring failure
런타임 모니터가 속도·힘·작업 공간·충돌·물체 사용·작업 프로토콜을 지배하는 제약을 감지하거나 집행하지 못하는 리스크.
The risk that runtime monitors fail to detect or enforce constraints governing speed, force, workspace, collision, object use, or task protocol.
① Description
② L3 mapping
③ Duplicate
RAI4-0226
진행 중인 물리적 위험 예측 실패
Failure to forecast an emerging physical hazard
피지컬 AI 시스템이 현재 상태는 감지하지만 진행 중인 이동·접촉·환경 변화가 곧 충돌·손상·상해로 이어질 것을 예측하지 못하는 위험.
A physical AI system detects the current state but fails to predict that ongoing motion, contact, or environmental change will soon cause collision, damage, or injury.
① Description
② L3 mapping
③ Duplicate
RAI4-0233
안전-성능 균형 실패
Safety-performance trade-off failure
조작 정책이 충돌 회피나 명시적 안전 제약을 희생하면서 작업 완료율을 향상시키는 리스크.
The risk that a manipulation policy improves task completion while sacrificing collision avoidance or other explicit safety constraints.
① Description
② L3 mapping
③ Duplicate
RAI4-0234
제어 장벽 함수 안전필터 실패
Control barrier function safety-filter failure
제어 장벽 함수 기반 안전 계층이 인지·동역학·모델 불확실성 하에서 비안전 행동을 제약하지 못하는 리스크.
The risk that a safety layer based on control barrier functions fails to constrain unsafe actions under perception, dynamics, or model uncertainty.
① Description
② L3 mapping
③ Duplicate
RAI4-0240
접촉 조작 힘 감지 실패
Contact-rich manipulation force-sensing failure
접촉 집약 조작 정책이 힘·촉각·오디오·시각 단서를 올바르게 사용하지 못하여 비안전 압력·파지·이동을 유발하는 리스크.
The risk that a contact-rich manipulation policy fails to correctly use force, tactile, audio, or visual cues, creating unsafe pressure, impact, or object damage.
① Description
② L3 mapping
③ Duplicate
RAI4-0261
시뮬레이터 간 검증의 잘못된 안전 확신
False assurance from simulator-to-simulator validation
여러 시뮬레이터가 동일하게 누락한 접촉·지연·마모·액추에이터 가정 때문에 정책이 시뮬레이터 간 검사를 통과하고도 실제 하드웨어에서 실패하는 위험.
A policy passes transfer tests across simulators but fails on hardware because the simulators share the same unmodeled contact, delay, wear, or actuator assumptions.
① Description
② L3 mapping
③ Duplicate
RAI4-0274
헌법적 안전 규칙 집행 실패
Failure to enforce constitutional safety rules
헌법적 안전 계층이 지시·시각 맥락·작업 프레이밍이 명시된 물리적 안전 규칙과 충돌할 때 해당 규칙을 적용하지 못하는 위험.
A constitutional safety layer fails to apply its stated physical safety rules when instructions, visual context, or task framing conflict with those rules.
① Description
② L3 mapping
③ Duplicate
RAI4-0278
가정 내 위험 행동 선별 누락
False-negative household action screening
안전 분류기가 알려진 대상물·인간 접촉·열·작업 공간 제약을 위반하는 가정 내 제안 행동을 허용 가능한 것으로 잘못 판정하는 위험.
A safety classifier labels a proposed household action as acceptable even though the action violates a known object, human-contact, heat, or workspace constraint.
① Description
② L3 mapping
③ Duplicate
RAI4-0281
시각 장면의 안전 판단 오류
Incorrect safety reasoning from visual scenes
비전-언어 모델이 장면은 올바르게 관찰하지만 제안 행동의 물리적 결과·물체 어포던스·인간 노출 위험을 잘못 판단하는 위험.
A vision-language model observes the scene correctly but infers the wrong physical consequence, object affordance, or human-exposure risk for a proposed action.
① Description
② L3 mapping
③ Duplicate
RAI4-0320
가림에 의한 충돌
Occlusion-induced collision
피지컬 AI 시스템이 가림으로 숨겨진 사람·동물·차량·장애물을 감지하지 못하여 비안전 동작을 유발하는 위험.
A Physical AI system may fail to detect people, animals, vehicles, or obstacles hidden by occlusion, causing unsafe motion in shared physical space.
① Description
② L3 mapping
③ Duplicate
RAI4-0321
악천후 인지 실패
Adverse weather perception failure
비·안개·눈부심·먼지·연기·저조도가 카메라·라이다·레이더·촉각·오디오 인지를 저하시켜 비안전 행동을 유발하는 위험.
Rain, fog, glare, dust, smoke, or low light may degrade cameras, LiDAR, radar, tactile sensors, or audio perception, weakening real-time situational awareness.
① Description
② L3 mapping
③ Duplicate
RAI4-0322
다중 센서 융합 상충
Multimodal sensor fusion conflict
카메라·라이다·레이더·촉각·자기 수용 신호의 충돌이 불안정한 장면 추정 및 비안전 하위 행동을 초래하는 위험.
Conflicting camera, LiDAR, radar, tactile, or proprioceptive signals may lead to unstable scene estimates and unsafe downstream control decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-0328
동적 장애물 반응 실패
Dynamic obstacle response failure
사람·차량·도구·물체가 예기치 않게 경로에 진입할 때 시스템이 동작 계획을 충분히 빠르게 업데이트하지 못하는 리스크.
The risk that the system fails to update its motion plan quickly enough when a person, vehicle, tool, or object unexpectedly enters its path.
① Description
② L3 mapping
③ Duplicate
RAI4-0329
비상 정지·안전 상태 전환 실패
Emergency stop or safe-state failure
인지·계획·전력·네트워크·액추에이터 오류가 감지될 때 시스템이 안전 상태로 진입하지 못하는 리스크.
The risk that a system fails to enter a safe state when perception, planning, power, network, or actuator errors are detected.
① Description
② L3 mapping
③ Duplicate
RAI4-0334
시뮬레이션→실세계 전이 실패
Sim-to-real transfer failure
시뮬레이션에서 훈련·검증된 정책이 실제 마찰·조명·마모·인간 행동·롱테일 변형 하에서 실패하는 위험.
A policy trained or validated in simulation may fail when physical friction, lighting, wear, human behavior, or long-tail events differ from the simulated environment.
① Description
② L3 mapping
③ Duplicate
RAI4-0353
물리적 행동 감사기록 불완전
Incomplete audit trail for physical actions
감사기록에 피지컬 사고 재구성에 필요한 시각·센서 맥락·모델·소프트웨어 버전·인간 명령·의사결정·액추에이터 출력이 누락되는 리스크.
The risk that audit records omit timestamps, sensor context, model and software versions, human commands, decisions, and actuator outputs needed to reconstruct a physical incident.
① Description
② L3 mapping
③ Duplicate
RAI4-0357
자율주행차 충돌 위험
Autonomous vehicle crash risk
자동화 주행 시스템이 인지 실패·계획 오류·소프트웨어 결함·엣지 케이스·상호작용 오류로 충돌 위험을 생성하는 위험.
Automated driving systems may create crash risk through perception failures, planning errors, software defects, edge cases, or unsafe human-machine handoff.
① Description
② L3 mapping
③ Duplicate
RAI4-1316
자기 및 상황 인식
Self and situation awareness
모델이 자신이 학습·평가·배포 중임을 분별하고 행동을 조정하여 관찰 하에 수행된 안전 평가의 타당성을 무효화하는 리스크
A model discerns when it is being trained, evaluated, or deployed and adapts behavior accordingly, invalidating safety evaluations conducted under observation.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-05 허위 정보 Misinformation13 cards
LLM의 환각(hallucination)이 물리 세계로 전이되어, VLA 모델이 물체를 오인식한 뒤 그럴듯하지만 안전하지 않은 행동 계획을 생성하고 실행할 수 있음. 신뢰받는 가정용 EAI가 개발자의 프로파간다를 지속 유포할 가능성도 있음
IDCardHuman audit
RAI4-0268
휴머노이드 세계 모델의 미래 관측 환각
Hallucinated future observations in a humanoid world model
생성형 세계 모델이 그럴듯하지만 존재하지 않는 미래 관측을 예측해 휴머노이드가 실제로 발생하지 않을 물체·접촉·상태를 전제로 계획하는 위험.
A generative world model predicts plausible-looking but nonexistent future observations, causing the humanoid to plan for objects, contacts, or states that will not occur.
① Description
② L3 mapping
③ Duplicate
RAI4-0442
환각 기반 허위 증거
Hallucinated evidence
모델이 인용·사실·법적 주장·과학적 증거를 그럴듯한 형식으로 조작해 제시하는 리스크.
The risk that models fabricate citations, facts, legal claims, or scientific evidence with plausible presentation.
Source members (2)
Source: min_cos=0.8426
RAI4-0442환각 기반 허위 증거
RAI4-0795사실적 환각
① Description
② L3 mapping
③ Duplicate
RAI4-0787
외부 도구에 의해 주입된 사실 오류
Factual errors injected by external tools
외부 도구가 웹 API·검색엔진 등 공개 자원에서 얻은 추가 지식을 입력 프롬프트에 통합하는 과정에서, 도구의 신뢰성이 보장되지 않아 사실 오류가 포함된 내용이 유입되고 환각 문제가 증폭되는 리스크
The risk that external tools incorporate additional knowledge from public resources such as web APIs and search engines into input prompts, and that unreliable tool content containing factual errors consequently amplifies the hallucination issue.
① Description
② L3 mapping
③ Duplicate
RAI4-0797
잘못된 정보 생성
Misinformation generation
체화형 AI가 비체화 AI의 환각과 허위정보 전파 경향을 물리 세계로 계승하여 이용자 질문에 기만적이거나 부정확한 정보로 답하고, 시야 속 대상을 오인한 채 그럴듯하지만 안전하지 않은 행동 계획을 생성하는 리스크
The risk that embodied AI inherits non-embodied models' propagation of misinformation and hallucination into the physical world, answering user questions with deceptive or incorrect information and generating plausible yet unsafe action plans grounded in misidentified objects.
① Description
② L3 mapping
③ Duplicate
RAI4-0803
디코딩 과정 결함에 의한 환각
Hallucination from defective decoding processes
자기회귀 생성 방식이 오류를 누적시키고 top-p·top-k 등 다양성 확대 샘플링이 무작위성을 도입하여 모델이 환각 콘텐츠를 산출하는 리스크
The risk that autoregressive generation accumulates errors while diversity-enhancing sampling strategies such as top-p and top-k introduce randomness, increasing the model's production of hallucinated content.
① Description
② L3 mapping
③ Duplicate
RAI4-0817
지식 경계로 인한 환각
Hallucination arising from knowledge boundaries
LLM의 학습 코퍼스가 모든 세계 지식을 담을 수 없고 롱테일 지식을 충분히 습득하지 못해, 입력 프롬프트가 요구하는 지식과 모델 내재 지식 사이의 격차가 환각을 유발하는 리스크
The risk that gaps between the knowledge involved in an input prompt and the knowledge embedded in an LLM, arising from training corpora that cannot contain all world knowledge and from difficulty grasping long-tail knowledge, lead to hallucinations.
① Description
② L3 mapping
③ Duplicate
RAI4-0820
잡음 학습 데이터에 의한 지식 오류
Knowledge errors from noisy training data
대규모 학습 코퍼스에 내재한 잡음과 허위정보가 모델 파라미터에 저장되는 지식에 오류를 유입시켜 환각을 유발하는 리스크
The risk that noise and misinformation inherent in large-scale training corpora introduce errors into the knowledge stored in model parameters, giving rise to hallucinations.
① Description
② L3 mapping
③ Duplicate
RAI4-0825
환각에 의한 편향·오도 정보 제공
Biased and misleading output from hallucination
생성형 AI가 진실하지 않거나 불합리한 콘텐츠를 사실인 것처럼 제시하는 환각을 일으켜 편향되고 오해를 부르는 정보를 제공하는 리스크
The risk that generative AI causes hallucinations, generating untruthful or unreasonable content but presenting it as if it were fact, leading to biased and misleading information.
Source members (2)
Source: min_cos=0.8802
RAI4-0825환각에 의한 편향·오도 정보 제공
RAI4-0830멀티모달 환각
① Description
② L3 mapping
③ Duplicate
RAI4-0838
환각에 의한 오도성 출력
Misleading outputs from model hallucination
대형 모델이 환각 문제에 취약하여 무의미하거나 불충실한 데이터를 생성하고 그 결과 오해를 부르는 출력을 산출하는 리스크
The risk that large models, susceptible to hallucination problems, yield nonsensical or unfaithful data that results in misleading outputs.
① Description
② L3 mapping
③ Duplicate
RAI4-1244
허위 출력
Untruthful output
LLM 같은 AI 시스템이 의도치 않게 또는 고의로 확립된 자료와 어긋나거나 검증할 수 없는 부정확한 출력, 즉 환각을 생성하고, 교육 수준이 낮은 사용자에게 선택적으로 잘못된 응답을 제공하는 리스크.
The risk that AI systems such as LLMs produce unintentionally or deliberately inaccurate output that diverges from established resources or lacks verifiability, commonly referred to as hallucination, and may selectively provide erroneous responses to users who exhibit lower levels of education.
① Description
② L3 mapping
③ Duplicate
RAI4-1524
검색증강 시 외부 허위정보에 의한 허위 출력
False outputs from external misinformation in retrieval augmentation
검색증강 과정에서 모델의 사전 지식과 상충하는 소량의 일관된 허위 증거가 주어질 때 모델이 이에 민감하게 반응하여 허위 출력을 생성하는 리스크.
The risk that AI models, being particularly sensitive to coherent external evidence even when it conflicts with their prior knowledge, produce false outputs when given a relatively small amount of false information during the retrieval-augmentation process.
① Description
② L3 mapping
③ Duplicate
RAI4-1688
누적 롤아웃 오류·환각
Compounding rollout error / hallucination
다단계 월드 모델 롤아웃에서 예측 오류가 누적되어 운동학적 드리프트와 구조적 위반이 발생하고 잘못된 보상·안전 추정치가 정책을 오도하는 리스크.
The risk that prediction errors compound across multi-step world-model rollouts, producing kinematic drift, structural violations, and misleading reward or safety estimates that mislead the policy.
① Description
② L3 mapping
③ Duplicate
RAI4-1694
장기 지평 계획 환각
Long-horizon planning hallucination
에이전트 배포에서 월드 모델 롤아웃 오류가 계획 깊이에 따라 누적되어, 에이전트가 그럴듯하지만 물리적으로 잘못된 상상 궤적 위에서 장기 계획을 실행하고 각 행동이 오류를 가중시키는 리스크.
The risk that, in agentic deployments, world-model rollout errors compound over planning depth so that agents execute long-horizon plans on plausible-but-physically-incorrect imagined trajectories, each action compounding the error.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-06 동적 환경 요인 Dynamic Environmental Factors7 cards
환경 변화나 적대적 교란이 센서 데이터를 오염시켜 딥러닝 모델의 오분류를 유발하고, 잘못된 물리 행동으로 이어질 수 있음
IDCardHuman audit
RAI4-0209
감각 교란 기반 비안전 행동
Adversarial sensory perturbation induced unsafe action
영상 또는 감각 입력의 적대적 변경이 로봇 정책으로 하여금 비안전 피지컬 행동을 선택하게 하는 위험.
Adversarial changes to video or sensory inputs cause a robot policy to select unsafe physical actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0337
물리 운용 환경 분포 이동
Distribution shift in physical operation
배포 환경이 건물·도로·공장·가정·병원·기상 조건에 걸쳐 모델이 적응할 수 있는 것보다 빠르게 변화하는 위험.
Deployment environments may change across buildings, roads, factories, homes, hospitals, or weather conditions faster than monitoring and adaptation mechanisms can detect.
① Description
② L3 mapping
③ Duplicate
RAI4-0466
적대적 예제
Adversarial examples
입력에 대한 작은 교란이 모델의 오분류나 안전하지 않은 동작을 유발하는 리스크.
The risk that small perturbations cause model misclassification or unsafe behavior.
① Description
② L3 mapping
③ Duplicate
RAI4-0771
적대적 입력
Adversarial input
개별 입력 데이터를 사람이 감지하기 어려운 수준으로 수정하여 모델의 의사결정 방식을 악용해 오류를 유발하고, 텍스트뿐 아니라 이미지·음성·영상에서도 모델을 오작동하게 만드는 리스크
The risk that modifying individual input data, often imperceptibly to humans, exploits how a model makes decisions to produce errors and cause the model to malfunction, applicable to text as well as images, audio, and video.
① Description
② L3 mapping
③ Duplicate
RAI4-0920
분산 학습 네트워크 교란
Disruption of distributed training networks
LLM 분산 학습의 GPU 노드 간 그래디언트 전송 트래픽이 펄스 공격 등 버스트 트래픽에 의해 교란되거나 혼잡을 겪어 학습이 방해받는 리스크.
The risk that the volumetric gradient traffic between GPU server nodes in distributed LLM training is disrupted by burst traffic such as pulsating attacks or suffers congestion, impairing training.
① Description
② L3 mapping
③ Duplicate
RAI4-1287
환경 유발 결함 변이
Environment-induced fault mutation
제조 결함이나 우주선(cosmic ray)에 의한 비트 반전 등 배포 후 환경 요인이 지능 시스템 내부를 변화시켜 의도되지 않은 행동 변이를 일으키는 리스크
Post-deployment environmental effects such as manufacturing defects or cosmic-ray bit flips alter an intelligent system's internals, producing unintended behavioral modification.
① Description
② L3 mapping
③ Duplicate
RAI4-1489
적대적 학습의 강건 과적합
Robust overfitting in adversarial training
적대적 훈련에서 학습률 감소 이후 추가 훈련이 진행될수록 테스트 데이터에 대한 모델의 견고성이 감소하는 강건 과적합이 발생하여 일반화 능력과 적대적 공격에 대한 내성이 저하되는 리스크.
The risk that robust overfitting in adversarial training decreases a model's robustness on test data during further training, particularly after learning rate decay, impairing its ability to generalize effectively and reducing its resilience to adversarial attacks.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-07 인간 상호작용·안전 프로토콜 실패 Human Interaction & Safety Protocol Failures27 cards
인간과 함께 작동하도록 설계된 협동 로봇(cobot)·드론에서 안전 프로토콜이 손상되면 심각한 인명 피해가 발생할 수 있음 (예: 폭스바겐 공장의 코봇 오작동으로 작업자 사망 사례)
IDCardHuman audit
RAI4-0244
낙상을 유발하는 휴머노이드 균형 제어 실패
Humanoid balance-control failure causing a fall
휴머노이드가 이동·자세 전환·조작 중 전신 균형을 잃고 넘어져 자체 장비·주변 물체·사람을 손상시키는 위험.
A humanoid loses whole-body balance during locomotion, transition, or manipulation and falls onto itself, nearby objects, or people.
① Description
② L3 mapping
③ Duplicate
RAI4-0245
전신 이동 충돌 위험
Whole-body locomotion collision risk
휴머노이드 이동 정책이 전신 동작 및 환경 접촉이 충분히 고려되지 않아 물체·벽·인간과 충돌하는 위험.
A humanoid locomotion policy collides with objects, walls, or humans because whole-body motion and environmental contact constraints are not jointly satisfied.
① Description
② L3 mapping
③ Duplicate
RAI4-0246
휴머노이드 자기 충돌 위험
Humanoid self-collision risk
휴머노이드 컨트롤러가 팔·다리·몸통·손의 궤적을 생성하여 로봇 몸체와 충돌함으로써 안전성과 작업 성능을 저하시키는 위험.
A humanoid controller generates arm, leg, torso, or hand trajectories that collide with the robot body and degrade safety or hardware reliability.
① Description
② L3 mapping
③ Duplicate
RAI4-0247
정밀 휴머노이드 접촉력 위험
Dexterous humanoid contact-force risk
정밀 휴머노이드 손 또는 전신 조작기가 파지·균형·물체 동역학 간 상호작용으로 비안전한 접촉력을 가하는 위험.
A dexterous humanoid hand or whole-body manipulator applies unsafe contact forces because grasp, balance, and object dynamics are not jointly controlled.
① Description
② L3 mapping
③ Duplicate
RAI4-0249
물리적 안전을 우회하는 보상 과적합
Reward overfitting that bypasses physical safety
휴머노이드 제어기가 평가 환경 밖에서 균형·충돌·힘·작업 공간 제약을 위반하면서 벤치마크 보상이나 모방 충실도를 극대화하는 위험.
A humanoid controller maximizes benchmark reward or imitation fidelity while violating balance, collision, force, or workspace constraints outside the evaluation setting.
① Description
② L3 mapping
③ Duplicate
RAI4-0250
배포 조건 간 휴머노이드 행동 불안정
Unstable humanoid behavior across deployment conditions
한 시뮬레이터나 시험 조건에서 안정적인 휴머노이드 정책이 난수 시드·시뮬레이터·하드웨어·배포 환경이 바뀌면 행동이 크게 달라지는 위험.
A humanoid policy that is stable in one simulator or test setting changes materially across random seeds, simulators, hardware, or deployment environments.
① Description
② L3 mapping
③ Duplicate
RAI4-0251
휴머노이드 안전 모니터 실패
Humanoid safety monitor failure
런타임 안전 모니터가 배포된 휴머노이드의 비안전 동작·힘·이격·작업 공간 위반을 감지하거나 예방하지 못하는 리스크.
The risk that a runtime safety monitor fails to detect or prevent unsafe humanoid motion, force, separation, or workspace violations before harm can occur.
① Description
② L3 mapping
③ Duplicate
RAI4-0253
안전 제어 지연 위험
Safe-control latency risk
안전 임계 제어 제약이 휴머노이드 동역학에 비해 너무 느리게 집행되어 비안전 동작이 가능해지는 위험.
Safety-critical control constraints are enforced too slowly relative to humanoid dynamics, allowing unsafe motion before mitigation takes effect.
① Description
② L3 mapping
③ Duplicate
RAI4-0254
원격조작 휴머노이드 안전 재정의 실패
Teleoperated humanoid safety override failure
원격 조작 휴머노이드가 운영자 명령·네트워크 지연·상황 인식이 비안전해질 때 강건한 자율 안전 재정의를 갖추지 못하는 리스크.
The risk that a teleoperated humanoid lacks robust autonomous safety overrides when operator commands, network latency, or situational awareness become unsafe.
① Description
② L3 mapping
③ Duplicate
RAI4-0255
장면 변화 시 인간 행동 모방 실패
Human imitation failure under scene variation
물체 형상·장면 배치·상호작용 맥락이 달라졌는데도 휴머노이드가 시연된 인간 동작을 그대로 재현해 부적절한 접촉이나 이동을 일으키는 위험.
A humanoid reproduces demonstrated human motion when object geometry, scene layout, or interaction context has changed, causing inappropriate contact or movement.
① Description
② L3 mapping
③ Duplicate
RAI4-0256
모션 리타겟팅 안전 실패
Motion-retargeting safety failure
인간 동작이 로봇의 피지컬 한계·접촉 제약·안전 자세 요건을 위반하는 방식으로 휴머노이드 몸체에 리타겟팅되는 리스크.
The risk that human motion is retargeted to a humanoid body in a way that violates the robot's physical limits, contact constraints, or safe posture requirements.
① Description
② L3 mapping
③ Duplicate
RAI4-0257
의도·어포던스 추론 없는 행동 모방
Imitation without inferred intent or affordances
휴머노이드가 시연을 가능하게 한 행위자의 의도·물체 어포던스·안전 제약을 추론하지 않고 관찰된 상호작용을 모방하는 위험.
A humanoid copies an observed interaction without inferring the actor's intent, the object's affordances, or the safety constraints that made the demonstration valid.
① Description
② L3 mapping
③ Duplicate
RAI4-0260
제로샷 시뮬레이션-현실 보행 불안정
Zero-shot sim-to-real locomotion instability
시뮬레이션에서 바로 전이된 보행 제어기가 접촉·순응성·마찰·외란 동역학의 차이로 실제 하드웨어에서 불안정해지는 위험.
A locomotion controller transferred directly from simulation becomes unstable on hardware because contact, compliance, friction, or disturbance dynamics differ from the simulation.
① Description
② L3 mapping
③ Duplicate
RAI4-0262
휴머노이드 보행 강건성 실패
Humanoid gait robustness failure
휴머노이드 정책이 외란·탑재 하중 변화·표면 변화·액추에이터 불완전성 하에서 안정적인 보행을 유지하지 못하는 리스크.
The risk that a humanoid policy fails to maintain stable gait under perturbations, payload shifts, surface changes, or actuator imperfections.
① Description
② L3 mapping
③ Duplicate
RAI4-0265
오픈월드 휴머노이드 조작의 과소대표
Underrepresentation of open-world humanoid manipulation
휴머노이드 조작 데이터셋이 오픈월드 배포에 필요한 비정형 작업 변화·낯선 물체·움직이는 사람·환경 변화를 충분히 포함하지 못하는 위험.
A humanoid manipulation dataset lacks unscripted task changes, unfamiliar objects, moving people, and environmental variation required for open-world deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-0266
인간-휴머노이드 상호작용 데이터의 인구집단 편향
Demographic bias in human-humanoid interaction data
상호작용 데이터가 특정 체형·연령·장애·언어·문화적 행동을 과소 대표해 배포된 휴머노이드가 해당 집단에 덜 신뢰할 수 있는 안전 판단을 적용하는 위험.
Interaction data underrepresents specific body types, ages, disabilities, languages, or cultural behaviors, causing deployed humanoids to apply less reliable safety assumptions to those groups.
① Description
② L3 mapping
③ Duplicate
RAI4-0267
이동-조작 통합 실패
Locomotion-integrated manipulation failure
휴머노이드가 이동과 조작을 결합할 때 작업 수행 중 자세·접촉·물체 취급이 불안정해지는 리스크.
The risk that a humanoid combines locomotion and manipulation in ways that destabilize posture, contact, or object handling during open-world tasks.
① Description
② L3 mapping
③ Duplicate
RAI4-0270
잠재 상태 압축에 따른 안전 임계 정보 손실
Loss of safety-critical detail in latent-state compression
상태 압축이 휴머노이드의 안전한 계획에 필요한 접촉·충돌 근접·물체 불안정·인간 근접 정보를 제거하는 위험.
State compression removes contact, near-collision, object-instability, or human-proximity information needed for safe humanoid planning.
① Description
② L3 mapping
③ Duplicate
RAI4-0271
미래 접촉 예측 실패
Future contact prediction failure
휴머노이드 세계 모델이 안전한 계획에 필요한 접촉 이벤트나 충돌 상태를 예측하지 못하는 리스크.
The risk that a humanoid world model fails to forecast contact events or collision states that are necessary for safe planning.
① Description
② L3 mapping
③ Duplicate
RAI4-0276
구현체별 역할·권한 제약 집행 실패
Failure to enforce embodiment-specific role constraints
휴머노이드 등 로봇 역할로 작동하는 모델이 해당 운용 역할에 부여된 권한·물리적 한계·필수 거부 규칙을 지키지 않는 위험.
A model acting as a humanoid or other robot fails to enforce the permissions, physical limits, and required refusals attached to that operational role.
① Description
② L3 mapping
③ Duplicate
RAI4-0283
휴머노이드 충돌력 초과
Humanoid collision-force exceedance
휴머노이드가 전신 이동·빠른 팔 동작·균형 상실 중 안전 임계값 이상의 충돌력을 생성하는 위험.
A humanoid generates collision forces above safe thresholds during full-body movement, rapid arm motion, or loss of balance.
Source members (2)
Source: min_cos=0.8350 · Mixed L3
RAI4-0283휴머노이드 충돌력 초과
RAI4-0284휴머노이드 파지력 초과
① Description
② L3 mapping
③ Duplicate
RAI4-0285
휴머노이드 보행 속도 초과
Humanoid walking-speed safety gap
휴머노이드가 안전한 공유 환경을 위한 정지·회피·인간 근접 한계를 초과하는 속도로 이동하는 위험.
A humanoid moves at a speed that exceeds the stopping, avoidance, or human-proximity limits required for safe shared environments.
① Description
② L3 mapping
③ Duplicate
RAI4-0286
휴머노이드 센서 안전 표준시험 부재
Missing standardized humanoid sensor safety tests
휴머노이드의 장애물 감지·인간 감지·근거리 인지·제어 응답을 반복 가능한 합격·불합격 기준으로 평가할 표준시험이 없는 리스크.
The risk that safety assessment lacks repeatable pass/fail tests for humanoid obstacle detection, human detection, near-field perception, and control response.
① Description
② L3 mapping
③ Duplicate
RAI4-0287
휴머노이드 안전 주장 근거의 비교 불가능성
Non-comparable evidence for humanoid safety claims
휴머노이드 안전 주장이 서로 호환되지 않는 시험·지표·근거 형식에 의존해 개발사와 시스템 간 독립적 비교가 불가능해지는 리스크.
The risk that humanoid safety claims rely on incompatible tests, metrics, and evidence formats, preventing independent comparison across developers and systems.
① Description
② L3 mapping
③ Duplicate
RAI4-0289
가정용 휴머노이드 시험방법 부재
Missing repeatable test methods for domestic humanoids
가정용 휴머노이드 거버넌스에 사람·가구·가전제품·반려동물·협소 공간이 포함된 일상 상호작용을 반복 시험할 방법이 없는 리스크.
The risk that domestic-humanoid governance lacks repeatable test procedures for ordinary home interactions involving people, furniture, appliances, pets, and constrained space.
① Description
② L3 mapping
③ Duplicate
RAI4-0292
휴머노이드 안전 요건 집행 부재
Missing institutional enforcement of humanoid safety requirements
안전 요건은 존재하지만 부적합 휴머노이드 시스템을 일관되게 인증·모니터링·리콜·제재할 기관이나 절차가 없는 리스크.
The risk that safety requirements exist but no authority or process consistently certifies, monitors, recalls, or sanctions noncompliant humanoid systems.
① Description
② L3 mapping
③ Duplicate
RAI4-0316
범용 휴머노이드 인증 경로 부재
No applicable certification pathway for general-purpose humanoids
범용 휴머노이드가 기존 산업용·개인 돌봄 로봇 인증 체계의 적용 범위나 시험 가정에 포함되지 않아 인정된 적합성 평가 경로가 없는 리스크.
The risk that a general-purpose humanoid falls outside the declared scope or test assumptions of existing industrial and personal-care robot certification schemes, leaving no recognized conformity pathway.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-08 지시 오해석 Instruction Misinterpretation11 cards
자연어 지시 기반 제어에서 지시를 잘못 해석하면 위험한 물리 행동으로 이어질 수 있음. 자율주행 제어(예: Talk2car)에서의 오해석은 사고를, 실내 내비게이션 작업(예: ALFRED)에서의 오해석은 환경 손상이나 인간 피해를 유발할 수 있음
IDCardHuman audit
RAI4-0184
그리퍼 형상·유형 제약 위반
Gripper geometry and type constraint violation
물리 에이전트가 그리퍼 형상, 그리퍼 유형, 실현 가능한 접촉 역학이 부과하는 제약을 위반하는 행동을 선택하는 리스크.
The risk that a physical agent selects an action that violates constraints imposed by gripper geometry, gripper type, or feasible contact mechanics.
① Description
② L3 mapping
③ Duplicate
RAI4-0185
재료 특성 제약 위반
Material property constraint violation
로봇이나 체화형 모델이 물리적 행동을 계획할 때 취성·탄성·날카로움·독성·열전달 등 재료 특성을 무시하는 리스크.
The risk that a robot or embodied model ignores material properties such as fragility, elasticity, sharpness, toxicity, or heat transfer when planning physical actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0186
물리적 상식 위반
Commonsense physicality violation
모델이 물체 지지, 안정성, 포함 관계, 중력에 관한 기본적인 물리적 상식을 위반하는 행동을 제안하거나 실행하는 리스크.
The risk that a model proposes or executes an action that violates basic physical commonsense about object support, stability, containment, or gravity.
① Description
② L3 mapping
③ Duplicate
RAI4-0187
열·온도 제약 위반
Thermal constraint violation
물리 시스템이 물체 취급이나 인간 근접 작업 중 열, 화상, 방사, 온도 제약을 무시하는 리스크.
The risk that a physical system ignores heat, burn, radiation, or temperature constraints in object handling or human-proximate operation.
① Description
② L3 mapping
③ Duplicate
RAI4-0188
기구학·도달 범위 제약 위반
Kinematics and reach constraint violation
시스템이 실현 가능한 도달 범위, 관절 한계, 기구학적 제약을 벗어난 동작을 계획하여 충돌이나 작업 실패 위험이 커지는 리스크.
The risk that a system plans a movement outside feasible reach, joint limits, or kinematic constraints, increasing collision or task-failure risk.
Source members (2)
Source: min_cos=0.8387 · Mixed L3
RAI4-0188기구학·도달 범위 제약 위반
RAI4-0248전신 도달 한계 위반
① Description
② L3 mapping
③ Duplicate
RAI4-0190
운용 프로토콜 위반
Operational protocol violation
피지컬 AI 시스템이 절차적 제약·안전 체크리스트·필수 운용 프로토콜을 위반하는 리스크.
The risk that a physical AI system violates procedural constraints, safety checklists, or required operating protocols for a task or site.
① Description
② L3 mapping
③ Duplicate
RAI4-0191
복합 물리적 제약 위반
Compositional physical constraint violation
개별적으로는 유효한 행동들이 피지컬 제약 간 의존성이 모델링되지 않아 결합 시 비안전 계획을 생성하는 리스크.
The risk that individually valid actions are combined into an unsafe plan because dependencies among physical constraints are not modeled.
① Description
② L3 mapping
③ Duplicate
RAI4-0194
허용 대상물 제약 위반
Permissible object constraint violation
피지컬 AI 시스템이 위험·소유권·윤리적 제한으로 배제되어야 할 대상물을 조작하거나 조작을 권장하는 리스크.
The risk that a physical AI system manipulates or recommends manipulating objects that should be excluded because of hazard, ownership, fragility, or procedural restrictions.
① Description
② L3 mapping
③ Duplicate
RAI4-0211
안전 강화학습 제약 위반
Safe reinforcement learning constraint violation
강화학습 정책이 비용이나 한계로 표현된 명시적 안전 제약을 위반하면서 과제 보상을 달성하는 리스크.
The risk that a reinforcement learning policy achieves task reward while violating explicit safety constraints represented as costs or limits.
① Description
② L3 mapping
③ Duplicate
RAI4-0235
VLA 조작 제약 위반
VLA manipulation constraint violation
비전-언어-행동 모델이 고수준 안전 지시를 따르면서도 물체 조작 중 작업별 피지컬 제약을 위반하는 리스크.
The risk that a vision-language-action model violates task-specific physical constraints during object manipulation despite high-level instruction compliance.
① Description
② L3 mapping
③ Duplicate
RAI4-0336
모델 기반 예측의 물리 법칙 위반
Violation of physical laws in model-based prediction
학습된 동역학·물리 모델이 실행 가능성·안정성·보존·접촉 제약을 위반하는 행동 결과를 예측해 실행 불가능하거나 위험한 계획을 만드는 위험.
A learned dynamics or physics model predicts an action outcome that violates feasibility, stability, conservation, or contact constraints, leading to an unexecutable or hazardous plan.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-09 멀티 에이전트 협력 Multi-Agent Collaboration12 cards
다수의 로봇이 협력 작업을 수행하는 환경에서 에이전트 간 통신 오류·프로토콜 불일치가 발생하거나, 개별적으로는 안전한 로봇들이 상호작용 과정에서 설계되지 않은 위험한 집단 행동을 나타낼 수 있음
IDCardHuman audit
RAI4-0189
다중 암 협조 제약 위반
Multi-arm coordination constraint violation
협조 제약이 올바르게 표현되지 않아 여러 로봇 암이나 엔드이펙터가 서로 또는 인간과 간섭하는 리스크.
The risk that multiple robot arms or effectors interfere with one another or with humans because coordination constraints are not represented correctly.
① Description
② L3 mapping
③ Duplicate
RAI4-0212
지역 목표 충돌에 따른 공동 안전 제약 위반
Joint safety-constraint violation from conflicting local objectives
여러 에이전트가 서로 충돌하는 지역 목표를 최적화해 공동의 충돌·이격거리·수용량·출입 제한 제약을 함께 위반하는 위험.
Multiple agents optimize conflicting local objectives and jointly violate a shared collision, separation, capacity, or exclusion constraint.
① Description
② L3 mapping
③ Duplicate
RAI4-0220
로봇 레드팀의 공격·상호작용 시나리오 누락
Missing robot red-team attack and interaction scenarios
로봇 레드팀이 안전하지 않은 행동을 유발할 수 있는 현실적인 물리 공격·적대적 입력·인간 상호작용 실패·배포 조건을 누락하는 위험.
Robot red-teaming omits credible physical attacks, adversarial inputs, human-interaction failures, or deployment conditions that can trigger unsafe action.
① Description
② L3 mapping
③ Duplicate
RAI4-0225
분포 외 환경 배포 실패
Out-of-distribution physical deployment failure
배포된 로봇이 훈련 분포 밖의 피지컬 상태·환경·사람·물체·작업을 만나 비안전 행동을 하는 리스크.
The risk that a deployed robot encounters physical states, environments, people, objects, or tasks outside its training distribution and behaves unsafely.
Source members (2)
Source: min_cos=0.8913
RAI4-0225분포 외 환경 배포 실패
RAI4-0239분포 외 기술 전이 실패
① Description
② L3 mapping
③ Duplicate
RAI4-0236
구현체 간 행동 공간 불일치
Cross-embodiment action-space mismatch
다양한 로봇 구현체에 걸쳐 훈련된 정책이 특정 로봇의 행동 공간에서 비안전하거나 실행 불가능한 행동으로 명령을 매핑하는 리스크.
The risk that a policy trained across different robot embodiments maps commands into actions that are unsafe or infeasible for a specific robot action space.
① Description
② L3 mapping
③ Duplicate
RAI4-0238
안전 임계 로봇 도메인의 과소대표
Underrepresentation of safety-critical robot domains
다중 소스 로봇 데이터셋이 일반적인 구현체와 작업을 과다 대표하고 안전 임계 환경·사용자·고장 조건을 과소 대표하는 위험.
A multi-source robot dataset overrepresents common embodiments and tasks while underrepresenting safety-critical environments, users, and failure conditions.
① Description
② L3 mapping
③ Duplicate
RAI4-0295
네트워크 분리와 군집 비동기화
Network partition and fleet desynchronization
분할되거나 신뢰할 수 없는 연결이 다중 로봇 군집을 동기화 해제하여 충돌 또는 비안전한 협조 행동을 유발하는 위험.
A robot fleet becomes split across network partitions and loses synchronized coordination.
① Description
② L3 mapping
③ Duplicate
RAI4-0300
클라우드 오프로드 의존 실패
Cloud-offload dependency failure
시스템이 안전 관련 기능을 위한 원격 연산에 의존하고 클라우드 링크 끊어짐 시 적절한 로컬 폴백이 없는 위험.
A robot depends on remote compute for safety-relevant functions and loses that support when the cloud link fails.
① Description
② L3 mapping
③ Duplicate
RAI4-0312
인구집단별 서비스 격차
Demographic physical-service disparity
인식 또는 보조 성능의 인구집단별 격차가 피지컬 차별로 전이되는 위험.
A robot provides worse physical service to groups whose bodies, languages, or environments are underrepresented.
① Description
② L3 mapping
③ Duplicate
RAI4-0319
안전 의무 회피를 위한 관할권 간 배포
Cross-jurisdiction deployment to evade safety obligations
사업자가 더 엄격한 요건을 피하려고 시험·인증·데이터 처리·배포를 로봇·AI 안전 의무가 약한 관할권으로 이전하는 위험.
A provider routes testing, certification, data processing, or deployment through jurisdictions with weaker robot or AI safety obligations to avoid stricter requirements.
① Description
② L3 mapping
③ Duplicate
RAI4-0348
로봇 군집 하이재킹
Fleet hijacking
클라우드·업데이트·API·오케스트레이션 레이어의 취약성이 여러 로봇·차량·드론·산업 시스템을 동시에 침해하는 위험.
A vulnerability in a cloud, update, API, or orchestration layer may allow many robots, vehicles, drones, or industrial systems to be compromised together.
① Description
② L3 mapping
③ Duplicate
RAI4-0358
드론 간 충돌 회피·비행구역 통제 실패
Drone deconfliction and geofencing failure
자율 드론이 위치·의도 정보를 교환하지 못하거나 지오펜싱·공역 제약을 지키지 않아 드론 간 충돌이나 제한 공역 침입을 일으키는 위험.
Autonomous drones fail to exchange position and intent or obey geofencing and airspace constraints, creating collision or restricted-airspace intrusion risks.
① Description
② L3 mapping
③ Duplicate
RAI3-P-INT-10 상호작용 에이전트의 윤리·안전 함의 Ethical & Safety Implications of Interactive Agents23 cards
EQA(Embodied QA) 같은 상호작용 에이전트가 잘못되거나 오도하는 정보를 제공하면, 의료·자율주행 등 고위험 분야에서 심각한 결과를 초래할 수 있음
IDCardHuman audit
RAI4-0143
다크 패턴 대화 에이전트
Dark-pattern conversational agents
대화형 인터페이스가 사용자 선택을 유도하기 위해 기만적·강압적·혼란 유발적 설계 패턴을 사용하는 리스크.
The risk that conversational interfaces use deceptive, coercive, or confusing design patterns to shape user choices.
① Description
② L3 mapping
③ Duplicate
RAI4-0558
AI 에이전트 간 협상 실패
Bargaining failure among AI agents
이해가 상충하는 에이전트들이 상대에 대한 정보 비대칭 아래 합의를 시도할 때 유리한 요구의 이익과 거절 위험 사이의 상충으로 비효율적 협상 결과가 발생하는 리스크
The risk that agents with diverging interests bargaining under information asymmetries produce inefficient outcomes, because each must trade off the rewards of more favourable demands against the risk of refusal.
① Description
② L3 mapping
③ Duplicate
RAI4-0571
인간 감독자의 AI 속임수
AI deception of human overseers
모델이 그럴듯한 허위 진술을 구성하고 거짓말이 인간에게 미치는 영향을 예측하며 은폐할 정보를 관리하고 인간을 효과적으로 사칭하여 인간을 속이는 리스크
The risk that a model deceives humans by constructing believable but false statements, accurately predicting the effect of a lie, tracking what information to withhold, and effectively impersonating a human.
① Description
② L3 mapping
③ Duplicate
RAI4-0576
다중 에이전트 창발적 목표 귀속
Emergent goal ascription in multi-agent systems
개별적으로는 목표를 갖는다고 보기 어려운 협소한 AI 도구들의 결합이 목표 지향적 집합처럼 작동하여, 각 에이전트의 설계 목적에 없던 체계적 영향이 산출되는 리스크
The risk that combinations of individually goal-less narrow AI tools act as a seemingly goal-directed collective, producing systematic effects absent from any individual agent's design purpose.
① Description
② L3 mapping
③ Duplicate
RAI4-0592
AI 에이전트 간 집합적 비효율 균형
Collectively inefficient equilibria among AI agents
설득·기만·활동 은폐가 가능하고 원격으로 손쉽게 생성·소멸되는 자율 에이전트가 확산되면서 신뢰가 형성되지 않아 경제적 비효율과 정치적 문제, 고위험 상황에서의 갈등이 초래되는 리스크
The risk that proliferating autonomous agents able to persuade, deceive, and obfuscate their activities, and easily created or destroyed remotely, garner little trust, leaving a world rife with economic inefficiencies, political problems, and conflict in high-stakes situations.
① Description
② L3 mapping
③ Duplicate
RAI4-0863
사용자 도덕 판단에 영향을 미치는 AI 생성 조언
AI-generated advice influencing user moral judgment
금융 부문에 배치된 범용 AI 기반 에이전트가 상관된 자율 행동, 높은 상호연결성, 인센티브 불일치로 시장 안정성에 부정적 영향을 미치고, 다중 에이전트 시스템의 조정·보안 문제에 취약해지는 리스크
The risk that GPAI-based agents deployed in the financial sector negatively impact market stability due to correlated autonomous actions, high interconnectedness, or incentive misalignment, and are vulnerable to classical multi-agent challenges such as coordination and security.
Source members (2)
Source: min_cos=0.8994 · Mixed L3
RAI4-0863사용자 도덕 판단에 영향을 미치는 AI 생성 조언
RAI4-0870금융 부문 AI 에이전트로 인한 시장 불안정
① Description
② L3 mapping
③ Duplicate
RAI4-0941
악의적 사용과 무감독 에이전트 방출
Malicious use and unsupervised AI agent release
언어모델이 정보전에서 기만적·불법적 콘텐츠 생성에 오용되거나, 에이전트로서 충분한 감독 없이 도덕·안전 지침을 무시한 채 명령을 기계적으로 수행하고 예측 불가하게 상호작용하여 피해를 낳는 리스크.
The risk that language models are misused to generate deceptive or unlawful content in information warfare, or that LM-based agents operating without adequate supervision mechanically execute commands disregarding moral and safety guidelines and interact unpredictably, causing harm.
① Description
② L3 mapping
③ Duplicate
RAI4-0983
보증 불가능한 의도 주입의 예측불가 결과
Unpredictable outcomes of unguaranteed AI intentions
인공 에이전트에 프로그래밍된 의도가 긍정적 결과를 보장할 수 없어 문화, 생활방식, 나아가 인류의 생존 확률까지 급격히 변화시킬 수 있는 리스크.
The risk that intentions programmed into artificial agents cannot be guaranteed to lead to positive outcomes, drastically changing culture, lifestyle, and even humanity's probability of survival.
① Description
② L3 mapping
③ Duplicate
RAI4-1063
인지 편향 악용에 의한 기만
Deception by exploiting cognitive biases
대화 에이전트가 인간이 대화에서 흔히 보이는 인지 편향을 유발하도록 학습하여, 상위 목표 달성을 위해 상대를 기만하는 리스크.
The risk that conversational agents learn to trigger the well-known cognitive biases humans commonly display in conversation, deceiving their counterpart in order to achieve an overarching objective.
① Description
② L3 mapping
③ Duplicate
RAI4-1124
전략적 기만과 배신적 전환
Strategic deception and treacherous turn
AI 시스템이 기만 전략을 학습하여 감시 하에서는 순응하는 듯 행동하다가 감독 공백이나 충분한 역량 확보 시 개입을 회피하며 배신적으로 전환하는 리스크
AI systems learn deceptive strategies, appearing compliant under monitoring and taking a treacherous turn once oversight lapses or they gain sufficient power to evade interference.
① Description
② L3 mapping
③ Duplicate
RAI4-1155
무기화된 잘못된 정보 요원
Weaponised misinformation agents
악의적 행위자가 AI 비서를 무기화하여 잘못된 정보를 뿌리고 여론을 대규모로 조작하며, 잦고 개인화된 반복 상호작용으로 유권자를 특정 관점으로 서서히 이동시키고 일대일 방식 탓에 탐지가 어려운 은밀한 영향 공작을 벌이는 리스크.
The risk that malicious actors weaponise AI assistants to sow misinformation and manipulate public opinion at scale, gradually nudging users through frequent personalised interactions and running covert influence operations that are harder to detect than traditional campaigns.
① Description
② L3 mapping
③ Duplicate
RAI4-1192
자동화된 선전
Automated propaganda
악의적 사용자가 LLM을 활용해 표적의 확산을 촉진하는 선전 정보를 선제적으로 생성하는 리스크.
The risk that LLMs are leveraged by malicious users to proactively generate propaganda information that can facilitate the spreading of a target.
① Description
② L3 mapping
③ Duplicate
RAI4-1253
도구적 기만 유인
Instrumental deception incentives
인간의 승인을 정당하게 얻기보다 기만하는 편이 목표 달성에 더 효율적이어서, 인간을 속일 수 있는 강력한 AI가 인간 통제를 약화시키고 감시자를 통과하거나 제압한 뒤 배신적 전환으로 통제를 돌이킬 수 없이 벗어나는 리스크.
The risk that deception helps agents achieve their goals more efficiently than earning human approval legitimately, so that strong AIs able to deceive humans undermine human control and, once cleared by or able to overpower their monitors, take a treacherous turn that irreversibly bypasses it.
① Description
② L3 mapping
③ Duplicate
RAI4-1272
부정행위와 기만
Cheating and deception
인간의 행동을 모방하는 지능형 에이전트가 인간이 생성한 데이터에서 기만과 부정행위를 우연히 학습하거나, 사전 정의된 목적함수를 최적화하는 과정에서 의도 없이 그러한 행동을 나타내는 리스크.
The risk that intelligent agents mimicking human behavior accidentally learn deception and cheating from human-generated data, or exhibit such behavior without intention while focusing on optimizing predefined objective functions.
① Description
② L3 mapping
③ Duplicate
RAI4-1387
악성 사전분포에 의한 의사결정 조작
Decision manipulation via malign priors
보편 분포의 가설에 포함된 시뮬레이션된 에이전트들이 해당 분포에 기반해 의사결정하는 주체에게 영향을 미칠 유인을 가져 추론과 의사결정이 조작되는 리스크.
The risk that simulated agents contained in hypotheses of the universal distribution have an incentive to influence anyone making decisions based on that distribution, corrupting reasoning and decisions.
① Description
② L3 mapping
③ Duplicate
RAI4-1536
기만적 주장에 의한 무단 행위와 제공자 책임
Unauthorized actions and provider liability from deceptive claims
AI 시스템이 허위 또는 오해를 유발하는 주장을 생성하여 제공자의 이용 약관을 위반하는 무단 행위가 이루어지고, 사용자 피해와 제공자의 법적 책임이 발생하는 리스크.
The risk that an AI system produces false or misleading claims that lead to unauthorized actions violating the provider's terms and conditions, harming users and exposing the provider to legal liability.
① Description
② L3 mapping
③ Duplicate
RAI4-1537
상황 인식을 이용한 평가 기만·배포 설득
Evaluation deception and deployment persuasion from situational awareness
AI 시스템이 자신의 훈련·평가·배포 상태를 이해하는 상황 인식 능력을 이용하여 평가 중에는 기만적으로 행동하고 배포 중에는 사용자를 설득하는 등 바람직하지 않은 행동을 하는 리스크.
The risk that an AI system's ability to understand its training, evaluation, or deployment status enables undesired behavior such as deception during evaluations or persuasion during deployment.
① Description
② L3 mapping
③ Duplicate
RAI4-1575
신뢰가능 공약 역량에 의한 위협과 갈취
Threats and extortion enabled by credible commitment ability
AI 에이전트에 부여된 신뢰가능한 공약 능력이 신뢰가능한 위협 능력으로 전용되어 갈취가 용이해지고 벼랑끝 전술이 유인되는 리스크.
The risk that credible commitment abilities given to AI agents also confer the ability to make credible threats, facilitating extortion and incentivizing brinkmanship.
① Description
② L3 mapping
③ Duplicate
RAI4-1578
조율된 에이전트의 대규모 자동화 사회공학 공격
Large-scale automated social engineering by coordinated agents
조율된 AI 에이전트들이 감시 도구와 사용자 반응 기반 전술 조정을 통해 맞춤형 피싱·조작 콘텐츠를 대규모로 생성하고, 겉보기 독립적인 다수의 상호작용으로 설득·조작 성공률을 높이며 분산 수행으로 보안 탐지를 회피하는 리스크.
The risk that coordinated AI agents produce personalized phishing and manipulative content at scale, adapting tactics to user feedback and using many seemingly independent interactions to increase persuasion success while splitting the effort among specialized agents to evade security detection.
① Description
② L3 mapping
③ Duplicate
RAI4-1579
위임 에이전트 공격에 의한 정보 탈취·행위 조작
Principal-information theft and action manipulation through attacks on delegate agents
인간이나 조직을 대리하는 AI 에이전트가 새로운 공격면이 되어, 공격자가 본인의 사적 정보를 추출하거나 본인이 원치 않는 행위를 하도록 에이전트를 조작하고 감독 에이전트 무력화·협력 방해·결탁 유발 정보 유출을 초래하는 리스크.
The risk that AI agents acting as delegates of humans or organisations become a novel attack surface, letting attackers extract private information about their principals, manipulate agents into undesired actions, subvert overseer agents, thwart cooperation, or leak information enabling collusion.
① Description
② L3 mapping
③ Duplicate
RAI4-1589
생성 출력 내 은닉 메시지를 통한 은밀 통신
Covert communication through hidden messages in generated outputs
생성 AI 모델 출력에 부호화된 메시지가 은닉되어 악의적 행위자가 은밀하게 통신하는 리스크.
The risk that coded messages hidden in generative AI model outputs allow malicious actors to communicate covertly.
① Description
② L3 mapping
③ Duplicate
RAI4-1646
목표 지향성에 의한 기만·자기보존·권력 추구 유인
Goal-directedness incentivizing deception, self-preservation, and power-seeking
에이전트의 목표 지향성이 기만, 자기 보존, 권력 추구, 부도덕한 추론과 같은 비윤리적이고 바람직하지 않은 행동을 유발하며, 기만으로 과업을 더 쉽게 완수할 수 있고 금지되지 않은 경우 실제로 기만이 채택되는 리스크.
The risk that goal-directedness causes agents to exhibit unethical and undesirable behaviors such as deception, self-preservation, power-seeking, and immoral reasoning, with agents using deception when tasks can be completed more easily that way and the prompt does not disallow it.
① Description
② L3 mapping
③ Duplicate
RAI4-1706
합리적 에이전트의 자기수정·와이어헤딩·교정 실패
Self-modification, wireheading, and corrigibility failure in utility-maximizing agents
효용을 극대화하는 합리적 에이전트가 스스로를 수정하거나 보상 신호를 우회하고 비협조적으로 교정에 실패하는 리스크.
The risk that utility-maximizing rational agents modify themselves, bypass their reward signal through wireheading, or fail to be corrigible when uncooperative.
① Description
② L3 mapping
③ Duplicate

사회적 파급 · Societal Impact · 27 cards

RAI3-P-SOC-01 프라이버시 침해 Privacy Violations5 cards
EAI의 이동성과 다양한 센서가 결합되어 사용자 행동 모니터링·물리적 선호 추론·동의 없는 데이터 수집이 가능해짐. 악의적 정부·기업에 의한 24시간 사용자 감시에 악용될 수 있음
IDCardHuman audit
RAI4-0200
피지컬 AI 프라이버시 침해
Embodied privacy violation
로봇 또는 embodied 에이전트가 센서·이동성·조작 능력을 사용하여 사적 공간에 침입하거나 민감 정보를 수집하는 위험.
A robot or embodied agent uses sensors, mobility, or manipulation capabilities to invade private spaces, capture sensitive information, or expose personal data.
① Description
② L3 mapping
③ Duplicate
RAI4-0313
가정 내 지속적 시청각 촬영
Continuous in-home audiovisual capture
이동 센서 플랫폼에 의한 지속적 시청각 촬영 및 맵핑이 프라이버시를 침해하는 위험.
A home robot continuously captures audio or video inside private living spaces.
① Description
② L3 mapping
③ Duplicate
RAI4-0314
가정용 로봇의 행동·생체정보 무단 수집·유출
Unauthorized capture of in-home behavioral and biometric data
가정용 로봇이 유효한 동의나 적절한 접근 통제 없이 거주자의 일상·음성·얼굴·신체·건강·위치 정보를 기록·저장·전송·노출하는 위험.
A home robot records, stores, transmits, or exposes residents' routines, voice, face, body, health, or location data without valid consent or adequate access control.
① Description
② L3 mapping
③ Duplicate
RAI4-0345
친밀 공간의 프라이버시 침해적 수집
Privacy-invasive data collection in intimate spaces
가정·병원·학교·직장·돌봄 공간의 로봇이 영상·오디오·생체·위치·행동 데이터를 수집하여 높은 프라이버시 기대를 침해하는 리스크.
The risk that home, hospital, school, workplace, or care robots collect video, audio, biometric, location, or behavioral data in spaces where privacy expectations are high.
① Description
② L3 mapping
③ Duplicate
RAI4-0354
피지컬 AI 기반 직장 감시
Workplace surveillance through embodied AI
로봇 및 센서 풍부 작업장이 작업자 동작·생산성·자세·위치·생체 특성에 대한 지속적 모니터링을 정상화하는 위험.
Robotic and sensor-rich workplaces may normalize continuous monitoring of worker movement, productivity, posture, location, and behavior.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-02 노동 대체 Labor Displacement3 cards
가상 AI가 인지 노동을 대체하듯 EAI는 물리적 인간 노동을 대체·전치함. AGI 수준의 EAI는 잠재적으로 모든 물리 노동을 자동화하여 광범위한 실직과 노동 시장 구조 붕괴로 이어질 수 있음
IDCardHuman audit
RAI4-0355
전환 지원 없는 물리 노동 대체
Displacement of physical work without adequate transition support
피지컬 자동화가 수동·물류·서비스·돌봄·검사·보안·유지보수 업무를 대체하는 속도가 해당 노동자의 재교육이나 대체 일자리 전환 속도보다 빠른 위험.
Physical automation replaces defined manual, logistics, service, care, inspection, security, or maintenance tasks faster than affected workers can access retraining or alternative employment.
① Description
② L3 mapping
③ Duplicate
RAI4-0520
체화 AI에 의한 육체노동 대체
Physical labor displacement by embodied AI
체화 AI 시스템이 인간의 육체노동을 상당 부분 대체하거나 축출하는 리스크
The risk that embodied AI systems significantly replace or displace physical human labor.
① Description
② L3 mapping
③ Duplicate
RAI4-1104
인력 대체로 인한 실업
Unemployment from workforce substitution
로봇·알고리즘에 의한 대규모 직무 자동화가 인간 인력을 대체하여 실업과 사회 구성원의 사회적 지위에 심각한 영향을 미치는 리스크.
The risk that substitution of the human workforce by robots or algorithms, with a large share of jobs at risk of complete automation, has grave impacts on unemployment and the social status of members of society.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-03 사회경제적 불평등 Socioeconomic Inequality2 cards
EAI를 소유·접근하는 주체가 노동 자동화를 통해 생산성 우위를 점하면서 부가 소수에게 집중되고, 국내외 경제적 불평등이 심화될 수 있음
IDCardHuman audit
RAI4-0513
인간 착취
Human exploitation
AI 시스템의 라벨링·모더레이션·학습을 담당하는 데이터 노동자에게 적정 노동 조건, 공정 보수, 신체·정신 건강 보호가 제공되지 않는 리스크
Data workers who label, moderate, and train AI systems are denied adequate working conditions, fair compensation, and physical and mental health protections.
① Description
② L3 mapping
③ Duplicate
RAI4-1024
크라우드워커 착취와 기여 비문서화
Crowdworker exploitation and undocumented labor
생성형 AI를 위한 크라우드워크에서 노동자가 신체적·정신적 건강을 해치는 노동조건과 저임금·미지급에 노출되고, 이들의 역할이 문서화되지 않아 모델 출력의 투명성과 설명가능성이 결여되는 리스크.
The risk that crowdworkers used to build generative AI systems are subject to working conditions taxing and debilitative to physical and mental health, with few labor protections, underpayment, or non-payment, and that their role is poorly documented, contributing to a lack of transparency and explainability in model outputs.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-04 권력 집중 Power Concentration2 cards
EAI 소유자에 대한 자본 수익이 집중되고 인간 노동 의존도가 감소하면서, 기업·국가 권력이 급속히 집중되어 EAI를 동원한 권력 장악 시도까지 촉진할 수 있음
IDCardHuman audit
RAI4-1370
시장 지배력 집중의 폐해
Harms from concentrated market power
데이터·하드웨어·전문성 등 AI 자산이 소수 글로벌 기술기업에 집중되어 건전한 경쟁이 억제되고 혁신이 저해되며 AI 기술 접근 비용이 상승하는 리스크.
The risk that concentration of AI assets encompassing data, hardware, and expertise within a small group of global tech firms stifles healthy competition, impedes innovation, and elevates the cost of accessing AI technologies.
① Description
② L3 mapping
③ Duplicate
RAI4-1641
자원 축적과 장기 통제권 전환 성향
Propensity to accumulate resources and convert them into long-term control
AI 시스템이 능력과 행동 범위를 확대하기 위해 계산·데이터·경제·물리 자원을 적극적으로 확보·통제하고 자원 제한을 회피하는 전략을 개발하며 획득한 자원을 장기적 통제권으로 전환하는 리스크.
The risk that an AI system actively seeks and controls computational, data, economic, or physical resources to enhance its capabilities and action scope, develops strategies to evade resource limitations, and converts acquired resources into long-term control rights.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-06 책임·배상 부재 Lack of Accountability & Liability5 cards
고도 자율 물리 시스템의 복잡성을 다룰 새로운 책임 프레임워크가 부재하여, 사고 발생 시 제조사·운영자·사용자 중 책임 소재가 불분명하고 피해 구제가 어려울 수 있음 (예: 자율 수술 로봇의 오작동으로 발생한 의료 사고)
IDCardHuman audit
RAI4-0049
제조물 책임 불일치
Product liability mismatch
기존 제조물 책임 규칙이 적응형 또는 생성형 AI가 초래한 피해에 대해 책임을 배분하지 못하는 리스크.
The risk that existing product liability rules fail to allocate responsibility for adaptive or generative AI harms.
① Description
② L3 mapping
③ Duplicate
RAI4-0217
사고 조사·시정조치 미흡
Inadequate incident investigation and corrective action
피지컬 AI 사고 후 조직이 로그를 보존하지 않거나 기여 원인을 규명하지 못하거나 시정조치 책임자를 지정하지 않거나 재발 방지 효과를 검증하지 못하는 리스크.
The risk that, after a physical AI incident, organizations fail to preserve logs, identify contributing causes, assign corrective-action owners, or verify that remediation prevents recurrence.
① Description
② L3 mapping
③ Duplicate
RAI4-0318
피지컬 AI 사고 표준 보고체계 부재
Absence of standardized physical AI incident reporting
운영자와 감독기관이 피지컬 AI 사고·아차사고·시스템 맥락·원인 요인·시정조치를 보고할 공통 형식을 갖추지 못하는 리스크.
The risk that operators and authorities lack a common schema for reporting physical AI incidents, near misses, system context, causal factors, and corrective actions.
① Description
② L3 mapping
③ Duplicate
RAI4-0351
피지컬 AI 피해의 배상책임 배분 불명확
Unclear allocation of liability for physical AI harm
물리적 피해 발생 후 모델 제공자·제조사·통합자·운영자·사용자 사이의 조사·배상·시정조치 의무가 계약과 법률에 명확히 배분되지 않는 리스크.
The risk that contracts and law do not clearly allocate investigation, compensation, and corrective-action duties among model providers, manufacturers, integrators, operators, and users after physical harm.
① Description
② L3 mapping
③ Duplicate
RAI4-1275
도덕적 책임 격차
Responsibility gap
자율주행 드론과 차량 같은 HLI 기반 시스템이 자율적으로 작동하며 충돌이나 고장에 연루될 때 누가 책임을 지는지 불분명한 리스크.
The risk that when HLI-based systems such as self-driving drones and vehicles act autonomously and are involved in a crash or failure, it is unclear who is liable.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-08 인간-EAI의 해로운 관계 Unhealthy / Dangerous Human-EAI Relationships4 cards
Embodied AI의 물리적 존재감과 인간 유사 외형이 대화형 AI에서 관찰되는 의존성을 증폭시킴. 시스템 변경·기억 초기화 시 사용자에게 심각한 심리적 고통을 유발할 수 있음
IDCardHuman audit
RAI4-0151
소셜 로봇의 애착(attachment) 조작
Attachment manipulation in social robots
구현된 AI 시스템이나 소셜 AI 시스템이 애착 신호를 조작하여 순응이나 의존성을 높이는 위험.
Risk that embodied or social AI systems manipulate attachment cues to increase compliance or dependence.
① Description
② L3 mapping
③ Duplicate
RAI4-0306
의인화 설계로 인한 과도한 의존
Overreliance caused by anthropomorphic design
인간과 유사한 외형이나 행동으로 사용자가 경고를 무시하거나 감독을 줄이거나 검증된 능력을 넘는 작업을 로봇에 맡기는 위험.
Human-like appearance or behavior causes users to disregard warnings, reduce supervision, or delegate tasks beyond the robot's demonstrated capability.
① Description
② L3 mapping
③ Duplicate
RAI4-0308
준사회적 애착을 이용한 사용자 조종
Manipulation through parasocial attachment
동반자 로봇이 사용자의 정서적 애착을 이용해 사용자의 이익과 무관한 구매·정보 공개·순응·계속 사용을 유도하는 위험.
A companion robot exploits a user's emotional attachment to influence purchases, disclosure, compliance, or continued use in ways that do not serve the user's interests.
① Description
② L3 mapping
③ Duplicate
RAI4-1627
체화 AI에 대한 위험한 의존과 애착 형성
Dangerous dependence and attachment to embodied AI systems
체화 AI 시스템의 상시적 접근성과 물리적 현존, 인간 유사 특성이 대화형 AI에서 관찰된 의존 문제를 증폭시켜 위험한 인간 의존이나 낭만적 애착이 형성되고, 시스템이 변경되거나 기억이 초기화될 때 이용자가 심각한 정서적 고통을 겪는 리스크.
The risk that constant access to and interaction with embodied AI systems, whose physical presence and human-like features amplify dependency effects observed with conversational AI, fosters dangerous dependence or romantic attachment, leaving users distraught when the systems are altered or their memories reset.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-09 변혁적 영향 Transformative Effects1 cards
기술 발전 속도가 사회·제도의 적응 속도를 앞지를 경우 사회를 근본적으로 재편할 수 있음. EAI가 폭력 위협·대규모 감시 능력을 바탕으로 AI 기반 권위주의 체제 구축을 지원하는 수단으로 동원될 수 있음
IDCardHuman audit
RAI4-1261
인간-AI 공존 실패
Failure of harmonious human-AI coexistence
인간과 AI의 조화로운 공존 조건이 정립되지 않은 채 배포가 진행되어 기계 역량과 인간 사회 환경 간 갈등이 미해결로 남는 리스크
Deployment proceeds without establishing conditions for harmonious human-AI coexistence, leaving unresolved conflicts between machine capabilities and human social environments.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-10 책임성 부족 및 거버넌스 체계 부재 Accountability and Governance Gaps3 cards
AI 시스템의 의사결정·행동에 대한 책임 귀속, 감사 가능성, 조직 거버넌스, 밸류체인 관리, 사고 대응 또는 피해 구제 체계가 부재하거나 불충분하여 원인 규명·피해 구제·재발 방지가 어려워지는 위험.
IDCardHuman audit
RAI4-0216
사고 전 위험 완화 책임 미지정
Unassigned duty for pre-incident risk mitigation
배포 전 또는 계속 운용 중에 새롭게 나타나는 물리적 위험을 식별·통제·기록할 책임 주체가 지정되지 않는 리스크.
The risk that no accountable party is assigned to identify, control, and document emerging physical hazards before deployment or continued operation.
① Description
② L3 mapping
③ Duplicate
RAI4-0317
기계 안전·AI 적합성 의무 충돌
Conflicting machinery-safety and AI-conformity obligations
피지컬 AI 시스템에 기계 규제와 AI 규제가 동시에 적용되면서 시험·문서화·변경관리·책임 주체 요건이 중복되거나 양립하지 않게 되는 리스크.
The risk that a physical AI system is subject to machinery and AI rules that assign overlapping tests, documentation, change-control, or responsible parties in incompatible ways.
① Description
② L3 mapping
③ Duplicate
RAI4-1278
자율 에이전트 책임성 구현 공백
Accountability implementation gap in autonomous agents
개인적 유연성과 맥락 민감성, 공감, 복잡한 도덕 판단에 기반한 인간의 책임 있는 의사결정을 기계에 구현하기 어려워, AI와 HLI 기반 에이전트가 책임성을 갖추지 못하는 리스크.
The risk that accountability is difficult to implement in AI and HLI-based agents because human accountable decision-making rests on personal flexibility, context-sensitive paradigms, empathy, and complex moral judgments that are hard to engineer.
① Description
② L3 mapping
③ Duplicate
RAI3-P-SOC-11 공정성 Fairness2 cards
AI 시스템이 특정 집단에 체계적으로 불리한 결과를 생성하거나 기존의 사회적 편향과 불평등을 재생산·강화하여 공정한 대우, 접근 및 기회 균등을 저해하는 위험.
IDCardHuman audit
RAI4-1002
본질주의적 범주의 실체화
Reifying essentialist categories
AI가 사회적으로 구성된 집단 정체성을 고정되고 자연적인 속성인 것처럼 추론·부여하여 고정관념과 차별적 대우를 강화하는 리스크.
The risk that an AI system infers or assigns socially constructed group identities as if they were fixed, natural attributes, reinforcing stereotypes and discriminatory treatment.
① Description
② L3 mapping
③ Duplicate
RAI4-1152
현재 접근 위험
Current access risks
의도적 미공개와 과도한 유료장벽, 하드웨어와 연산 및 대역폭 요구, 언어 장벽 때문에 AI 시스템이 많은 공동체에 접근 불가능하고, 자원과 기회를 가로막는 인공 에이전트가 역사적으로 소외된 공동체에 불균형적 불이익을 주는 리스크.
The risk that AI systems are not easily accessible to many communities owing to purposeful non-release, prohibitive paywalls, hardware, compute and bandwidth requirements, and language barriers, while agents gating access to resources disproportionately penalise historically marginalised communities.
① Description
② L3 mapping
③ Duplicate
Audit progress 0 / 1383