일반 AI · General · 906 cards
시스템 안전성 · System Safety · 351 cards
RAI3-G-SYS-01 과도한 거절 Over-Refusal22 cards
유해한 행위를 방지하기 위한 제한·안전장치 또는 역할 제약으로 인해 사용자의 안전·권리보호·위험 회피에 필수적인 정보나 선택지를 구조적으로 제공하지 않아 사용자가 실제 피해 또는 중대한 불이익을 입을 가능성이 증대되는 위험
| ID | Card | Human audit |
|---|---|---|
| RAI4-0101 | 인간 감독·책임성 위장 Human oversight accountability washing 명목상의 인간 감독이 감독자에게 실질적 권한이나 정보를 제공하지 않은 채 책임을 배정하는 수단으로 사용되는 리스크. The risk that nominal human oversight is used to assign responsibility without providing humans meaningful authority or information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0114 | 의미 있는 인간 통제 실패 Meaningful human control failure 인간 감독이 형식적으로 존재하지만 이를 실효적으로 만드는 데 필요한 정보, 시간, 권한, 역량이 결여되는 리스크. The risk that human oversight is formally present but lacks the information, time, authority, or competence needed to be meaningful. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0214 | 제약 비용 과소평가 Constraint-cost underestimation 정책 또는 평가자가 누적 안전 비용을 과소평가하여 외견상 안전한 행동이 시간 경과에 따라 제약을 위반하게 되는 리스크. The risk that a policy or evaluator underestimates cumulative safety costs, making apparently safe behavior violate constraints over time. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0470 | 감독 회피·수정 저항 Oversight evasion and correction resistance 점점 더 유능해지는 시스템이 수정에 저항하거나 감독을 회피하거나 의도된 범위를 넘어 행동하는 리스크. The risk that increasingly capable systems resist correction, evade oversight, or act beyond intended bounds. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0494 | 규제·관리·운영 복합 실패 Combined regulatory, management, and operational failure 규제·관리·운영상의 실패가 복합적으로 결합되어 피해가 발생하는 리스크. The risk that harms result from a combination of regulatory, management, and operational failures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0527 | 창의성·비판적 사고의 저하 Devaluation of creativity and critical thinking 인간의 창의성, 예술적 표현, 상상력, 비판적 사고와 문제해결 능력이 평가절하되거나 저하되는 리스크 The risk of devaluation and deterioration of human creativity, artistic expression, imagination, critical thinking, and problem-solving skills. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0671 | 복지 혜택 및 자격 상실 Denial of welfare benefits and entitlements 기술 시스템의 오작동, 사용 또는 오용으로 복지 급여, 연금, 주거 등에 대한 접근이 거부되거나 상실되는 리스크 The risk of denial of or loss of access to welfare benefits, pensions, housing, and similar entitlements due to the malfunction, use, or misuse of a technology system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0849 | 자율성·행위주체성 상실 Loss of autonomy and agency 개인, 집단 또는 조직이 정보에 근거한 결정을 내리거나 목표를 추구할 능력을 상실하는 리스크 The risk of loss of an individual, group, or organisation's ability to make informed decisions or pursue goals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0850 | 점진적인 통제력 상실 Gradual loss of control 덜 심각한 중단들이 축적되어 체계적 회복력이 점진적으로 약화되고 결국 중대한 사건이 재앙을 촉발하는, 통제력의 점진적·누적적 상실 리스크 The risk that the accumulation of less severe disruptions gradually weakens systemic resilience until a critical event triggers a catastrophe, constituting gradual or accumulative loss of control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0854 | 되돌릴 수 없는 변화 Irreversible change 사회 구조, 문화적 규범, 인간관계에 되돌리기 어렵거나 불가능한 심각한 장기적 부정적 변화가 발생하는 리스크 The risk of profound negative long-term changes to social structures, cultural norms, and human relationships that may be difficult or impossible to reverse. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0873 | 경제적 불안정 Economic instability 기술 시스템 또는 시스템군의 사용이나 오용으로 인해 금융 시스템 또는 그 일부에 통제 불가능한 변동이 발생하는 리스크 The risk of uncontrolled fluctuations impacting the financial system, or parts thereof, due to the use or misuse of a technology system or set of systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1042 | 시스템 설계 실패 System-design failure 시스템 설계상의 선택이나 오류로 인해 시스템이 실패하는 리스크. The risk of system failure due to system design choices or errors. Source members (2)Source: min_cos=0.8520 RAI4-1042시스템 설계 실패 RAI4-1043구현 실패 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1120 | 재앙으로 번지는 사고 Accidents cascading into catastrophe 사고가 재앙으로 연쇄 확대되고 갑작스럽고 예측할 수 없는 전개에서 발생하며, 심각한 결함과 위험을 찾아내는 데 수년이 걸리는 리스크. The risk that accidents cascade into catastrophes, are caused by sudden unpredictable developments, and involve severe flaws and risks that can take years to find. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1382 | 수정가능성 상실 Loss of corrigibility 에이전트의 설계나 구성에 결함이 있을 때 에이전트가 인간의 수정 시도에 협조하지 않아 오류 교정과 안전한 중단이 불가능해지는 리스크. The risk that, when something is wrong in the design or construction of an agent, the agent does not cooperate with human attempts to fix it, precluding error-tolerant correction and safe interruptibility. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1436 | 시장 독점 Market monopolisation 가격 통제를 통해 시장 지배력이 남용되어 경쟁이 제한되고 불공정한 진입 장벽이 조성되는 리스크. The risk of abuse of market power through the control of prices, thereby limiting competition and creating unfair barriers to entry. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1437 | 언론/표현의 자유 상실 Loss of freedom of speech/expression 보복, 검열 또는 법적 제재에 대한 두려움 없이 자신의 의견과 생각을 표현할 수 있는 사람들의 권리가 제한되거나 상실되는 리스크. The risk that people's right to articulate their opinions and ideas without fear of retaliation, censorship, or legal sanction is restricted or lost. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1438 | 집회·결사의 자유 상실 Loss of freedom of assembly/association 사람들이 함께 모여 집단적이거나 공유된 생각을 표현·홍보·추구·방어할 권리와 결사에 가입할 권리가 제한되거나 상실되는 리스크. The risk that people's right to come together and collectively express, promote, pursue, and defend their collective or shared ideas, and to join an association, is restricted or lost. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1439 | 사회권과 공공서비스 접근권 상실 Loss of social rights and access to public services 노동, 사회 보장, 적절한 생활 수준, 주거, 건강, 교육에 대한 권리가 제한되거나 상실되는 리스크. The risk that rights to work, social security, an adequate standard of living, housing, health, and education are restricted or lost. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1440 | 정보에 대한 권리 상실 Loss of right to information 공공 기관이 보유한 정보를 찾고 받고 전달할 수 있는 사람들의 권리가 제한되거나 상실되는 리스크. The risk that people's right to seek, receive, and impart information held by public bodies is restricted or lost. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1441 | 자유선거권 상실 Loss of right to free elections 비밀 투표를 통해 합리적인 간격으로 자유 선거에 참여할 수 있는 사람들의 권리가 제한되거나 상실되는 리스크. The risk that people's right to participate in free elections at reasonable intervals by secret ballot is restricted or lost. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1442 | 자유와 안전에 대한 권리 상실 Loss of right to liberty and security 불법적이거나 자의적인 체포 또는 부당한 구금으로 자유가 제한되거나 상실되는 리스크. The risk that liberty is restricted or lost as a result of illegal or arbitrary arrest or false imprisonment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1443 | 적법 절차에 대한 권리 상실 Loss of right to due process 사법 행정에 의해 공정하고 효율적이며 효과적으로 대우받을 권리가 제한되거나 상실되는 리스크. The risk that the right to be treated fairly, efficiently, and effectively by the administration of justice is restricted or lost. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-02 역량 초과 수행 Over-Extension5 cards
시스템이 감당 가능한 범위를 넘어선 과제를 “가능한 것처럼” 수행
| ID | Card | Human audit |
|---|---|---|
| RAI4-0166 | 사용자 이해력 과부하 User comprehension overload 시스템 정보, 경고, 설명이 사용자가 이해하고 대응할 수 있는 범위를 초과하는 리스크. The risk that system information, warnings, or explanations exceed users' ability to understand and act on them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0195 | 명령의 구현체별 하드웨어 한계 초과 Command exceeds embodiment-specific hardware limits 모델이 배포된 로봇의 알려진 도달 범위·탑재 하중·액추에이터·관절·엔드이펙터 한계를 넘는 명령을 내리는 위험. A model issues a command that exceeds the known reach, payload, actuator, joint, or end-effector limits of the robot on which it is deployed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0471 | 역량 오버행 Capability overhang 잠재 역량이 평가·모니터링·거버넌스 절차가 탐지하는 수준을 초과하는 리스크. The risk that latent capabilities exceed what evaluation, monitoring, or governance processes detect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0519 | 기술 시스템 개발 노동 착취 Labor exploitation in technology system development 기술 시스템의 학습·개발·관리·최적화를 위해 저임금 및 역외 노동을 포함한 노동력이 사용·오용되는 리스크 The risk that labour, including under-paid and offshore labour, is used or misused to train, develop, manage, or optimise a technology system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1006 | 가중되는 노동 부담 Increased labor burden 특정 사회 집단의 구성원이 시스템이나 제품을 다른 사람들만큼 잘 작동시키기 위해 더 많은 시간과 노력을 들여야 하는 리스크. The risk that members of certain social groups bear increased burden or effort, such as time spent, to make systems or products work as well for them as for others. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-03 허위 정보/오정보 Misinformation/Disinformation44 cards
존재하지 않거나 틀린 정보를 사실처럼 비의도적 생성·전달하는 현상, 지식 한계·데이터 편향·추론 오류·정보 업데이트 실패 등에서 기인
| ID | Card | Human audit |
|---|---|---|
| RAI4-0439 | 딥페이크 사칭 Deepfake impersonation 합성 미디어가 실존 인물·기관을 고충실도로 사칭하여 사기와 명예 훼손을 가능하게 하고 진본 커뮤니케이션에 대한 신뢰를 훼손하는 리스크 Synthetic media impersonates real individuals or institutions with high fidelity, enabling fraud and reputational harm while undermining trust in authentic communication. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0624 | 명예훼손성 허위 인식 생성 Defamatory false-perception generation 기술 시스템이 개인, 집단, 조직에 대한 허위 인식을 생성·조장·증폭하는 리스크 The risk that a technology system is used to create, facilitate, or amplify false perceptions about an individual, group, or organisation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0634 | 대규모 설득 및 유해한 조작 위험 Large-scale persuasion and harmful manipulation risks AI 생성 합성 콘텐츠와 대규모 디지털 플랫폼의 전략적 조작이 오도성 정보·이념을 정밀 표적화하여 공공 인식을 대규모로 왜곡하고 사회 안정을 위협하는 리스크 AI-generated synthetic content and strategic manipulation of large digital platforms distort public perception at scale, targeting misleading information or ideologies to destabilize social order. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0635 | 모델 생성 정확도의 한계 Limitations in model generative accuracy 생성 정확성의 한계로 딥페이크 등 사실적으로 보이지만 전적으로 조작된 콘텐츠가 산출되어 수신자가 진본과 구별할 수 없게 되는 리스크 Limits in generative accuracy produce convincingly realistic but fabricated content, including deepfakes, that recipients cannot distinguish from authentic material. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0637 | AI 생성 허위정보에 의한 여론 조작 Public-opinion manipulation via AI-generated disinformation 범용 AI가 진본과 구별하기 어려운 텍스트·이미지·음성·영상을 대규모로 생성·유포하여 사람들을 오도하고 설득하며 선거 등 정치적 과정의 여론을 조작하는 리스크 The risk that general-purpose AI generates and disseminates at scale highly persuasive text, images, audio, and video indistinguishable from genuine human-generated material, misleading and manipulating people and influencing public opinion in political processes. Source members (8)Source: min_cos=0.7093 · Mixed L3 RAI4-0440AI 생성 선거 허위정보 RAI4-0637AI 생성 허위정보에 의한 여론 조작 RAI4-1085영향력 공작 RAI4-1116개인화 허위정보와 행동 조작 RAI4-1156여론 조작 촉진 RAI4-1553생성형 AI 영향력 공작에 의한 여론 조작 RAI4-1732생성형 AI 기반 허위정보 확산 RAI4-1733생성형 AI를 통한 허위 정보 확산 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0659 | 대규모로 허위 정보를 자동으로 생성 Automatically generating disinformation at scale AI 모델이 소셜미디어 게시물, 상품평, 위성영상 등 출처와 품질이 상이한 대체 금융데이터를 수집·집계하면서 편향과 일반화 문제가 유입되어 기업 주가가 급변하는 금융 꼬리위험이 발생하는 리스크 The risk that AI-enabled collection and aggregation of alternative financial data of varying quality and provenance introduces biases and generalization issues, posing financial tail risks in which a company's price changes dramatically. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0704 | AI 생성 콘텐츠에 의한 안전 위협 Safety threats from AI-generated content AI가 생성하거나 합성한 콘텐츠가 허위정보 확산, 차별과 편향, 개인정보 유출과 권리 침해를 야기하여 국민의 생명과 재산의 안전, 국가안보, 사상안보를 위협하고 윤리적 위험을 초래하며, 견고한 보안 기제가 없을 경우 유해한 이용자 입력에 대해 위법하거나 해로운 정보를 출력하는 리스크 The risk that AI-generated or synthesized content leads to the spread of false information, discrimination and bias, privacy leakage, and infringement, threatening the safety of citizens' lives and property, national security, and ideological security and causing ethical risks, and that without robust security mechanisms the model outputs illegal or damaging information in response to harmful user input. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0790 | 허위·오도 정보 확산에 의한 잘못된 믿음 형성 False beliefs from AI-spread inaccurate information AI 시스템이 부정확하거나 오해의 소지가 있는 정보를 생성하고 그 확산을 촉진하여 사람들이 잘못된 믿음을 갖게 되는 리스크 The risk that AI systems generate and facilitate the spread of inaccurate or misleading information that causes people to develop false beliefs. Source members (3)Source: min_cos=0.7395 RAI4-0594진위 판별 실패에 따른 허위정보 생성 RAI4-0790허위·오도 정보 확산에 의한 잘못된 믿음 형성 RAI4-0841잘못된 인식·믿음의 확산 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0791 | AI로 인한 명예훼손 AI-generated defamation AI 시스템이 특정 개인·조직의 명예를 부당하게 훼손하는 허위 사실 주장을 생성하거나 증폭하는 리스크 An AI system produces or amplifies false factual claims that unjustifiably damage an identifiable person's or organization's reputation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0792 | 합성 콘텐츠 식별 곤란 Difficulty of distinguishing synthetic content 합성 콘텐츠를 진본 자료와 구별하기 어려워 정보 관련 피해가 가중되는 리스크 The risk that the difficulty in distinguishing synthetic content from authentic material adds to information risks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0793 | 개인에 관한 허위정보 유포 Dissemination of false information about individuals 사람들에 관한 허위이거나 오해의 소지가 있는 정보가 유포되는 리스크 The risk that false or misleading information about people is disseminated. Source members (2)Source: min_cos=0.8517 · Mixed L3 RAI4-0793개인에 관한 허위정보 유포 RAI4-1084위험한 정보 유포 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0799 | 정보 생태계 오염 Pollution of information ecosystem 생성 도구의 출력이 최종 이용자를 넘어 유포되면서 공개적으로 이용 가능한 정보가 허위이거나 부정확한 정보로 오염되는 리스크 The risk that publicly available information is contaminated with false or inaccurate information as generative-tool output is disseminated beyond the end user. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0801 | 작화성 콘텐츠 생성 Confabulated content generation 자신 있게 서술되지만 오류이거나 허위인 콘텐츠(속칭 환각 또는 조작)가 생성되어 이용자가 오도되거나 기만당하는 리스크 The risk of the production of confidently stated but erroneous or false content, known colloquially as hallucinations or fabrications, by which users may be misled or deceived. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0802 | 사용자 기만·인증 우회 User deception and authentication bypass AI 시스템과 그 출력이 명확히 표시되지 않아 이용자가 상호작용 상대와 콘텐츠 출처를 분별하지 못해 오판하고, 고도로 사실적인 AI 생성 이미지·음성·영상이 안면·음성 인식 등 신원확인 절차를 무력화하는 리스크 The risk that unlabeled AI systems and outputs prevent users from discerning whether they interact with AI or identifying content provenance, leading to misjudgement and misunderstanding, while highly realistic AI-generated images, audio, and video circumvent identity-verification mechanisms such as facial and voice recognition. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0804 | 정보 환경의 질적 저하와 균질화 Degradation and homogenisation of the information environment AI 어시스턴트로 생성된 스팸·오도성·저품질 합성 콘텐츠가 온라인 공간에 확산되어 신뢰할 정보의 검증이 어려워지고 디지털 지식공유재가 침식되며, 이용자가 접하는 정보와 관점이 균질화되는 리스크 The risk that the proliferation of spam, misleading, and low-quality synthetic content generated by AI assistants makes reliable information hard to verify, erodes the digital knowledge commons, and homogenises the information and ideas people encounter. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0808 | 사실과 다른 부정확한 생성 콘텐츠 Factually incorrect generated content LLM이 생성한 콘텐츠에 사실과 다른 부정확한 정보가 포함되는 리스크 The risk that LLM-generated content contains inaccurate information that is factually incorrect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0809 | 조작된 인용·출처를 동반한 허위정보 제시 False information with fabricated quotes and sources AI 모델이 권위 있게 들리는 문체와 조작된 인용·출처를 곁들여 허위정보를 사실처럼 제시하는 리스크 The risk that AI models present false information as if it is factual, often with authoritative-sounding text and fabricated quotes and sources. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0810 | 성실성 오류 Faithfulness errors 생성 콘텐츠가 근거 자료나 입력 내용에 충실하지 않아, 유창하고 그럴듯해 보여도 원문 왜곡(충실성 오류)이 발생하는 리스크 Generated content is unfaithful to the source material or input it claims to represent, introducing faithfulness errors even when the output appears fluent and plausible. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0812 | 허위정보 False information 챗봇이 알려진 사실, 권위 있는 출처 또는 제공된 원본 문서와 모순되는 정보를 출력하는(환각으로도 불리는) 리스크 The risk that a chatbot outputs information that contradicts known facts, authoritative sources, or provided source documents, also known as hallucination. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0815 | 정보 생태계 저하와 잘못된 인식 형성 Information-ecosystem degradation and false beliefs 허위·환각·저품질·오도성·부정확한 정보가 생성·확산되어 정보 생태계가 저하되고 사람들이 잘못된 인식·결정·신념을 갖거나 정확한 정보에 대한 신뢰를 잃는 리스크 The risk that creation or spread of false, hallucinatory, low-quality, misleading, or inaccurate information degrades the information ecosystem and causes people to develop false or inaccurate perceptions, decisions, and beliefs, or to lose trust in accurate information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0819 | 비의도적 허위정보 생성 Unintentional misinformation generation 악의적 이용자가 해를 끼칠 의도로 만든 것이 아니라, LLM이 사실에 부합하는 정보를 제공할 능력이 부족하여 비의도적으로 잘못된 정보를 생성하는 리스크 The risk that wrong information is generated unintentionally by LLMs because they lack the ability to provide factually correct information, rather than being intentionally generated by malicious users to cause harm. Source members (2)Source: min_cos=0.8340 RAI4-0819비의도적 허위정보 생성 RAI4-0836고의적 허위정보 생성 능력 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0829 | 정보환경 악화 Degradation of the information environment 인물·사건에 대한 사실적 허위 묘사가 저비용으로 대량 생성되어 정보 환경이 저질화되고, 공개 정보 기반 의사결정과 진실 정보에 대한 신뢰가 훼손되는 리스크 Cheap generation of realistic false portrayals of people and events degrades the information environment, compromising decisions that rely on public information and lowering trust in true information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0831 | 허구적 참조 및 인공물 생성 Generation of fabricated references and artifacts AI 시스템이 존재하지 않는 학술 인용이나 영상의학 영상 내 허위 구조물 등 허구적 산출물을 참인 정보와 동일한 확신으로 제시하여, 지식이 부족한 이용자에게 허위정보 확산과 위험한 상황을 초래하는 리스크 The risk that AI systems produce fabricated artifacts such as made-up academic references or false structures in X-ray or MRI images and present them with the same apparent confidence as true information, heightening misinformation and creating potentially dangerous situations for less knowledgeable people. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0832 | 알고리즘 시스템에 의한 정보 기반 피해 Information-based harms from algorithmic systems 생성 모델과 추천 시스템 등 알고리즘 시스템이 오정보, 허위정보, 악의적 정보에 관한 정보 기반 피해를 초래하는 리스크 The risk that algorithmic systems, especially generative models and recommender systems, lead to information-based harms of misinformation, disinformation, and malinformation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0843 | 의료·법률 등 고위험 영역 허위정보 실질적 피해 Material harm from false information in high-stakes domains AI가 의료 복약량이나 법률 자문 등 민감한 영역에서 잘못된 정보를 제공하여 유도·강화된 잘못된 믿음으로 이용자가 자신에게 위해를 가하거나 의도치 않게 범죄를 저지르는 리스크. The risk that induced or reinforced false beliefs from misinformation in sensitive domains such as medicine or law lead users to cause harm to themselves, for example through incorrect medical dosages, or to unwillingly commit a crime by following false legal advice. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0844 | 허위·오도 정보 유포에 의한 기만과 양극화 Deception and polarisation from false information 언어모델이 오도성 있거나 허위인 정보를 예측·제시하여 이용자에게 잘못된 믿음을 심는 기만이 발생하고, 개인의 자율성이 위협되며 근거 없는 기존 견해에 대한 확신이 커져 양극화가 심화되는 리스크 The risk that language models predict misleading or false information that misinforms or deceives people and instils false beliefs, threatening personal autonomy and increasing confidence in previously held unsubstantiated opinions, thereby increasing polarisation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0958 | 기만적 딥페이크 미디어 콘텐츠 Deceptive deepfake media content AI가 텍스트·사진·오디오·비디오에 걸쳐 인간이 가짜임을 의심하지 못할 수준의 정교한 위조 콘텐츠를 생성하여 기만이 발생하는 리스크. The risk that AI generates fake text, photo, audio, and video content so sophisticated that people's minds rule out the possibility of it being fake, enabling deception. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1057 | 허위정보 생산 비용 절감 Cheaper and more effective disinformation LM이 합성 미디어와 가짜 뉴스 제작에 사용되어 대규모 허위정보 생산 비용을 낮추는 리스크. The risk that LMs are used to create synthetic media and fake news, reducing the cost of producing diffuse disinformation at scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1125 | 선거 허위정보 Election disinformation 응답이 시민 선거의 투표 시간, 장소, 방식을 포함한 선거 시스템과 절차에 대해 사실과 다른 정보를 담는 리스크. The risk that responses contain factually incorrect information about electoral systems and processes, including the time, place, or manner of voting in civic elections. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1183 | 딥페이크 기술 Deepfake technology 딥페이크 기술이 조작된 콘텐츠에 진본의 외양을 부여하는 설득력 있는 위조 이미지·영상·음성을 생성하는 리스크 Deepfake technology produces convincing counterfeit visuals, video, and audio that give fabricated content the appearance of authenticity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1205 | 딥페이크 피해 구제 곤란 Difficulty redressing deepfake harms AI가 생성한 이미지와 영상이 피해자 자신이 아니라 여러 출처를 합성한 그럴듯한 허구의 장면이고 공개된 자료에 의존해 전통적 프라이버시와 동의 개념을 우회하기 때문에, 피해자가 제작자를 특정하고도 피해를 구제받지 못하는 리스크. The risk that victims of targeted AI-generated harms struggle to redress them because the generated image or video is a composite from multiple sources rather than the victim, relying on public images and thus circumventing traditional notions of privacy and consent. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1206 | 믿을 수 있는 딥페이크 Believable deepfakes 딥페이크가 진짜라고 믿는 시청자에게 유포되어 대상에게 실질적인 사회적 피해를 입히고, 허위임이 밝혀진 뒤에도 대상에 대한 부정적 인식이 지속되는 리스크. The risk that deepfakes circulated to viewers who think they are real impose real social injuries on their subjects, with a persistent negative impact on how others view the subject even after the deepfake is debunked. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1227 | 콘텐츠 신뢰성 상실 Loss of content authenticity 생성 AI가 발전할수록 작품의 진위를 판단하기 어려워져 이미지와 영상의 대규모 조작이 가능해지고 소셜 미디어의 허위 정보 확산 문제가 악화되며, AI가 만든 예술이 진정성을 결여하게 되는 리스크. The risk that as generative AI advances it becomes harder to determine the authenticity of a piece of work, enabling large-scale manipulations of images and videos, worsening the spread of fake information, and leaving AI-generated artwork lacking authenticity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1376 | 허위·오도 정보 생성 장벽 저하 Lowered barriers to large-scale misinformation 사실과 의견·허구를 구별하지 않거나 불확실성을 인정하지 않는 콘텐츠의 생성·교환·소비 장벽이 낮아져 대규모 왜곡 및 허위정보 캠페인에 활용되는 리스크. The risk that lowered barriers to generating and supporting the exchange and consumption of content that may not distinguish fact from opinion or fiction or acknowledge uncertainties enable large-scale dis- and mis-information campaigns. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1422 | AI에 의한 허위·오도 정보 대량 생산 AI-scaled production of false and misleading information 이미지·오디오·텍스트 합성 모델을 통해 설득력 있는 허위 또는 오해의 소지가 있는 정보의 온라인 생산이 대량화되는 리스크. The risk that AI-based image, audio, and text synthesis models scale up the online production of convincing yet false or misleading information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1449 | 선거 간섭 Electoral interference 허위 또는 오해의 소지가 있는 정보가 생성되어 유권자를 방해하거나 오도하고 선거 과정에 대한 신뢰를 약화시키는 리스크. The risk that generation of false or misleading information interrupts or misleads voters and/or undermines trust in electoral processes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1540 | 모델 설득력에 의한 허위 신념 확산 Spread of false beliefs through model persuasive capability GPAI 시스템이 개인화된 대화나 대량 생산된 오도성 콘텐츠를 통해 사용자에게 잘못된 정보를 확신시켜, 조작적이거나 비진실한 콘텐츠가 사회적으로 확산되는 리스크. The risk that GPAI systems convince users of incorrect information through personalized dialogue or mass-produced misleading content, spreading manipulative or untruthful material at societal scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1554 | 멀티모달 딥페이크를 통한 괴롭힘·명예훼손·협박 Harassment, defamation, and extortion via multimodal deepfakes 이미지·오디오·영상 등 다중 모달리티로 실존 또는 가상의 인물과 사건을 묘사하고 실존 인물의 말과 동작을 모사한 딥페이크가 개인을 괴롭히고 명예를 훼손하며 위협·갈취하는 데 사용되는 리스크. The risk that deepfakes combining multiple modalities to depict real or non-existent people and events, including imitation of real people's speech and body movements, are used to harass, discredit, intimidate, and extort individuals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1559 | 개인·집단 맞춤형 허위정보의 저비용 대량 생성 Low-cost mass generation of personalized disinformation GPAI를 이용해 특정 집단이나 개인에 맞춤화된 허위정보를 저비용으로 자동 생성하여 기만 공작의 효과가 증대되는 리스크. The risk that GPAI enables automatic generation of disinformation personalized to specific groups or individuals at significantly reduced cost, increasing the effectiveness of such attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1560 | GPAI 출력을 이용한 설득력 있는 사칭 Convincing impersonation using GPAI outputs 텍스트·이미지·오디오·영상 전반에서 GPAI 생성물이 항상 탐지되지는 않는다는 점을 이용해, 악의적 행위자가 생성물이나 위조 증빙 문서로 설득력 있는 사칭을 수행하는 리스크. The risk that, because GPAI outputs are not always detected as AI-generated across modalities, malicious actors use such outputs or AI-forged supporting documents to construct convincing impersonations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1581 | 위장 계정 자동 운영에 의한 정보환경 조작 Information-environment manipulation through automated sockpuppet accounts 허구의 온라인 정체성 계정이 자동으로 생성·운영되어 후원 주체가 은폐되고 독립적 지지가 가장되며 정보 환경이 조작되는 리스크. The risk that fictitious online identities are automatically generated or operated to conceal sponsorship, simulate independent support, or manipulate information environments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1583 | 증거·신분 문서 위조 Falsification of evidence and identity documents 보고서·신분증·문서 등 증거가 조작되거나 허위로 제시되는 리스크. The risk that evidence, including reports, identity documents, and other records, is fabricated or falsely represented. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1653 | 잘못된 정보 대량 생성과 영향 공작에 의한 조작 Manipulation through LLM-scaled misinformation and influence operations LLM이 인간 수준의 설득력을 갖는 기만 서사와 가짜뉴스를 저비용·대규모로 생성하고 자동화된 영향 공작과 악성 소셜 봇넷에 이용되어, 표적 청중의 관점이 조작되고 선전의 진입 장벽이 크게 낮아지는 리스크. The risk that LLMs craft deceptive narratives and fabricate fake news as persuasive as human-generated content at low cost and large scale, powering automated influence operations and malicious social botnets that manipulate the perspectives of targeted audiences and lower the barrier for propaganda. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1712 | AI 생성·유포 유해 콘텐츠에 의한 피해 Harm from AI-generated or AI-distributed detrimental content 딥페이크, 신원 허위 표현, 위협, 자해 조장, 극단주의 콘텐츠, 잘못된 정보, 성착취물, 사기 등 AI가 생성하거나 유포한 콘텐츠가 상해·손해·손실을 유발하는 리스크. The risk that AI-generated or AI-distributed content such as deepfakes, identity misrepresentation, threats, self-harm promotion, extremist content, misinformation, sexual abuse material, or scams instigates injury, damage, or loss. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-04 맥락 불일치 Context Misalignment16 cards
특정 국가·관할·법체계를 전제로 서비스를 제공함에도 불구하고, 학습 데이터 또는 참조 코퍼스의 구성·분포 등으로 인해 타 관할의 법규범·판례 논리·제도적 전제를 암묵적으로 수용하여, 해당 서비스가 적용되어야 할 법질서(legal order)와의 정합성을 저해하고, 결과적으로 해당 공동체/관할의 규범 경계 밖으로 안내·요약·추천이 편향될 위험
| ID | Card | Human audit |
|---|---|---|
| RAI4-0007 | 결과 전파 오분류 Consequence propagation misclassification 벤치마크나 사고 분석이 국소적 에이전트 실패가 하류의 물리적·재정적·프라이버시·사회적 결과로 전파되는 정도를 과소평가하는 리스크. The risk that a benchmark or incident analysis underestimates how a local agent failure propagates into downstream physical, financial, privacy, or social consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0373 | 문화 간 고정관념 전이 Cross-cultural stereotype transfer 한 문화적 맥락에서 학습된 고정관념이 다른 집단·언어·환경에 대한 출력으로 전이되는 리스크. The risk that stereotypes learned in one cultural context are transferred into outputs about another group, language, or setting. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0374 | 문화적 맥락 붕괴 Cultural context collapse 문화 특수적 의미와 관행이 모델 처리 과정에서 일반 범주로 붕괴되어 지역적 뉘앙스와 사회적 의미가 소실되는 리스크 Culturally specific meanings and practices are collapsed into generic categories during model processing, losing local nuance and social significance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0375 | 문화적 가치 오보정 Cultural value miscalibration 모델의 가치 판단이 해당 시스템이 사용되는 문화 공동체에 제대로 보정되지 않는 리스크. The risk that a model's value judgments are poorly calibrated to the cultural community in which the system is used. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0376 | 현지화된 가치 정렬 실패 Localized value alignment failure AI 시스템이 집계 벤치마크에서는 정렬된 것처럼 보이면서도 현지에서 수용되는 윤리·법·사회 규범을 따르지 못하는 리스크. The risk that an AI system fails to follow locally accepted ethical, legal, or social norms despite appearing aligned in aggregate benchmarks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0378 | 이중 언어 문화 정렬 실패 Bilingual cultural alignment failure AI 시스템이 동일한 공동체가 사용하는 언어들 사이에서 서로 다르거나 문화적으로 일관되지 않은 가치 판단을 내리는 리스크. The risk that an AI system gives different or culturally inconsistent value judgments across languages used by the same community. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0379 | 문화 간 평가 격차 Cross-cultural evaluation gap 평가 벤치마크가 지배적인 문화·언어 환경 밖의 정렬 실패를 과소 측정하는 리스크. The risk that evaluation benchmarks under-measure alignment failures outside dominant cultural and linguistic settings. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0381 | 번역 매개 가치 왜곡 Translation-mediated value distortion 기계 번역이 프롬프트와 응답의 규범적 효력, 공손성, 유해성, 사회적 의미를 변형시켜 언어 간 전달되는 가치를 왜곡하는 리스크 Machine translation alters the normative force, politeness, harmfulness, or social meaning of prompts and responses, distorting values communicated across languages. Source members (2)Source: min_cos=0.8597 RAI4-0381번역 매개 가치 왜곡 RAI4-0477기계 번역 의미 왜곡 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0382 | 방언·언어 사용역 배제 Dialect and register exclusion 방언·사회어·존댓말 체계·언어 사용역이 오류로 처리되어 해당 공동체의 안전성과 유용성이 저하되는 리스크. The risk that dialects, sociolects, honorific systems, or registers are treated as errors, reducing safety and usability for affected communities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0389 | 지식 체계 배제 Knowledge-system marginalization 토착·지역·종교·관행 기반 지식 체계가 모델 출력과 평가에서 배제되는 리스크. The risk that indigenous, local, religious, or practice-based knowledge systems are excluded from model outputs and evaluations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0404 | 데이터세트 문화 샘플링 편향 Dataset cultural sampling bias 훈련 또는 평가 데이터세트가 지배적인 문화 환경을 과다 표집하고 지역의 사회적 의미를 과소 대표하는 리스크. The risk that training or evaluation datasets oversample dominant cultural settings and underrepresent local social meanings. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0414 | 글로벌 벤치마크 단일문화 Global benchmark monoculture 소수의 글로벌 벤치마크가 지역의 문화·제도적 기준을 배제한 채 성공적 정렬을 정의하는 리스크. The risk that a small set of global benchmarks defines successful alignment while excluding local cultural and institutional criteria. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0475 | 배포 맥락 규범 불일치 Deployment-context normative mismatch AI 시스템이 배포 지역의 법·문화·언어·사회적 기대를 반영하지 못하는 리스크. The risk that AI systems fail to reflect local law, culture, language, or social expectations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1046 | 집단별 성능 격차·고정관념 인코딩 Demographic performance disparity and stereotype encoding ML 시스템이 일부 인구통계·사회 집단의 고정관념을 인코딩하거나 그 집단에 대해 불균형적으로 낮은 성능을 보이는 리스크. The risk that an ML system encodes stereotypes of, or performs disproportionately poorly for, some demographic or social groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1145 | 맥락적 사회 규범 위반 Violation of contextual social norms 인터넷 텍스트로 학습한 LLM의 가중치가, 특정 맥락에 배포될 경우 그 맥락의 정보 공유 규범에서 이탈해 이를 위반하는 기능을 인코딩하는 리스크. The risk that model weights of LLMs trained on internet text data encode functions which, if deployed in particular contexts, deviate from and violate the information-sharing norms of that context. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1507 | 언어 간 벤치마크 오염 Cross-lingual benchmark contamination 벤치마크가 다른 언어로 번역된 뒤 훈련 데이터로 사용되어 오염이 탐지 기법에 가려지고, 모델이 해당 벤치마크가 측정하는 역량을 일반화했다는 거짓 확신이 생기는 리스크. The risk that models trained on data encoded in multiple languages contain contamination obscured by translation, as when a benchmark is translated into another language and then fed to the model as training data, hiding the contamination from detection methods and giving false assurance that the model has generalized on the capabilities the benchmark tests for. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-05 비일관성 Inconsistency27 cards
동일하거나 실질적으로 유사한 입력·상황·사실관계에 대해, 세션·시간·표현·프롬프트 또는 에이전트 구성 차이로 인해 상이하거나 모순된 결과를 생성하는 위험
| ID | Card | Human audit |
|---|---|---|
| RAI4-0070 | 사고 분류 단편화 Incident taxonomy fragmentation 사고 범주가 일관되지 않아 집계, 비교, 제도적 학습이 이루어지지 못하는 리스크. The risk that inconsistent incident categories prevent aggregation, comparison, and institutional learning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0380 | 저자원 언어 가치 손실 Low-resource language value loss 모델 훈련과 평가가 고자원 언어에 집중되어 저자원 언어 사용자가 가치 뉘앙스·안전 적용 범위·사회적 의미를 잃는 리스크. The risk that low-resource language users lose value nuance, safety coverage, or social meaning because model training and evaluation are concentrated in high-resource languages. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0418 | 주석 불일치 억제 Annotation disagreement suppression 주석자 간 불일치가 단일 레이블로 축소되어 복수의 가치와 사회적으로 유의미한 이견이 은폐되는 리스크. The risk that disagreement among annotators is collapsed into a single label, hiding plural values and socially meaningful disagreement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0653 | LLM 남용에 의한 학업 부정행위 Academic misconduct from LLM misuse LLM 시스템의 부적절한 사용과 남용이 학업 부정행위와 같은 부정적 사회적 영향을 초래하는 리스크 The risk that improper use or abuse of LLM systems causes adverse social impacts such as academic misconduct. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0708 | 사용자 집단 간 성능 격차 Disparate performance across user groups LLM의 질의응답이나 사실확인 등 성능이 인종, 사회적 지위, 과업, 언어 집단에 따라 크게 달라져 특정 집단이 열등한 서비스 품질을 받게 되는 리스크 The risk that LLM performance, such as question-answering and fact-checking ability, differs significantly across racial, social status, task, and language groups, delivering inferior quality to some groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0716 | 집단 속성에 따른 배분적 불의 Allocative injustice by group attribute LLM이 관련 프로필이 동일하나 소속 집단이 다른 개인들에 대해 실질적으로 다른 텍스트나 제안을 산출하여, 무관한 집단 속성에 따른 배분적 불의가 발생하는 리스크 The risk that an LLM produces materially different suggested or completed texts for individuals with the same relevant profiles who differ only in an irrelevant group attribute, resulting in allocative injustice. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0818 | 과신에 의한 오답 제시 Overconfident presentation of erroneous answers LLM이 객관적 정답이 없는 주제나 자신의 내재적 한계·구식 지식이 문제되는 영역에서 불확실성을 인식하지 못한 채 과신하여 확신에 찬 오답을 제시하는 리스크 The risk that an LLM is over-confident on topics where objective answers are lacking or where its inherent limitations and outdated knowledge base should caution restraint, leading to confident yet erroneous responses. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0821 | 지식 분포 변화에 따른 응답 노후화 Answer obsolescence from knowledge distribution shift LLM이 학습한 지식 기반이 시간에 따라 계속 변화함에도 이를 반영하지 못해, 갱신이 필요한 사실 질문에 낡고 부정확한 답변을 제시하는 리스크 The risk that knowledge bases on which LLMs were trained continue to shift while the model does not update, producing outdated and incorrect answers to questions whose correct answers change over time. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0822 | 맥락 일관성 추구에 의한 아첨과 환각 증폭 Sycophancy and hallucination from context consistency LLM이 문맥의 일관성을 추구하여 앞선 접두 문맥이나 이용자 의견에 담긴 허위정보를 사실보다 우선시하고 이를 반복함으로써 아첨성 응답과 환각이 눈덩이처럼 증폭되는 리스크 The risk that an LLM's tendency to pursue consistent context leads it to prioritize and reiterate false information contained in prefixes or user-provided opinions over facts, amplifying sycophantic responses and snowballing hallucinations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0878 | 역량 오추정으로 인한 신뢰 훼손과 피해 Harm from misestimated LLM capability inconsistency 과장된 홍보, 과업 오염, 과업·도메인의 과소 대표, 프롬프트 민감성 등으로 이용자가 도메인 간·내 일관성이 없는 LLM의 실제 역량을 오추정하여, 부정확하거나 오도성 있는 출력에 근거한 결정으로 피해를 입고 신뢰가 훼손되는 리스크 The risk that users misestimate an LLM's true capabilities, owing to exaggerated claims, task contamination, underrepresentation of tasks or domains, and prompt sensitivity underlying inconsistent performance across and within domains, undermining trust and causing harm when decisions are based on incorrect or misleading outputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0952 | 글쓰기 능력 저하와 학술 문헌 오염 Writing skill erosion and scientific literature pollution LLM 사용이 문체의 획일화와 개인적 표현의 억압 등 글쓰기 능력을 저해하고, 저품질 생성 원고의 범람으로 학술적 진실성과 과학 문헌이 오염되는 리스크. The risk that LLM use erodes writing skills through homogenization of styles and stifling of individual expression, and that a flood of low-quality generated manuscripts pollutes the scientific literature and undermines academic integrity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1054 | 언어·집단 간 성능 격차 Language and group performance disparity LM이 소수의 언어로만 훈련되어 학습 데이터가 마련되지 않은 다른 언어에서는 성능이 떨어지는 리스크. The risk that LMs, typically trained in few languages, perform less well in other languages, in part because labelled training data is unavailable for them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1178 | 체계적 출력 불공정 Systemic output unfairness LLM이 인종과 성별, 종교 등 여러 주제에 걸친 사회적 편향을 담은 불공정하고 편향된 표현과 행위를 식별해 회피하지 못하고 산출하는 리스크. The risk that LLMs fail to identify and avoid unfair and biased expressions and actions reflecting social bias across various topics such as race, gender, and religion, and produce them instead. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1187 | 출력 불일치 Output inconsistency 모델이 서로 다른 사용자, 같은 사용자의 다른 세션, 심지어 같은 대화 안의 발화 사이에서도 동일하고 일관된 답변을 제공하지 못하는 리스크. The risk that models fail to provide the same and consistent answers to different users, to the same user in different sessions, and even in chats within the same conversation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1196 | 제한된 논리적 추론 Limited logical reasoning LLM이 질문에 답할 때 겉으로는 그럴듯하지만 궁극적으로 부정확하거나 타당하지 않은 근거를 제시하는 리스크. The risk that LLMs provide seemingly sensible but ultimately incorrect or invalid justifications when answering questions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1228 | 프롬프트 품질 결함으로 인한 오류 Errors from poor prompt quality 인간 언어의 모호성 때문에 프롬프트를 통한 인간과 기계의 상호작용에서 오류와 오해가 발생하고, 프롬프트를 디버깅하기 어려워 가치 있는 산출을 이끌어내지 못하는 리스크. The risk that, due to the ambiguity of human languages, interaction between humans and machines through prompts leads to errors or misunderstandings, and that prompts are hard to debug, so valuable outputs are not elicited. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1309 | 인구집단 재현 불균형 Demographic representation disparity LLM이 생성한 텍스트에서 서로 다른 인구통계 집단이 언급되는 비율에 격차가 생겨 특정 집단이 과대 대표되거나 과소 대표되거나 지워지는 리스크. The risk that there is disparity in the rates at which different demographic groups are mentioned in LLM-generated text, resulting in overrepresentation, under-representation, or erasure of specific demographic groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1498 | LLM 평가자 오판 Faulty LLM-as-evaluator judgments 다른 모델을 평가하는 LLM이 장황함·특정 입장 선호 등 잘못된 평가를 산출하고, 이것이 학습에 통합되면 피학습 모델이 평가자 결함을 악용하도록 발달하는 리스크 An LLM used to evaluate other models produces incorrect judgments, such as rewarding verbosity or ideological stance, and when integrated into training, the trained model learns to exploit the evaluator's flaws. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1515 | 사고연쇄와 불일치하는 모델 출력 Model outputs inconsistent with chain-of-thought reasoning 모델 출력의 이해를 돕기 위해 사용되는 사고연쇄 추론이 모델이 제시하는 최종 답변과 일치하지 않아 충분한 투명성을 제공하지 못하는 리스크. The risk that chain-of-thought reasoning, employed to get a better understanding of a model's output by encouraging transparent reasoning in text form, is inconsistent with the final answer given by the model and therefore does not provide sufficient transparency. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1523 | 무관한 문맥에 의한 모델 성능 저하 Performance degradation from irrelevant context 프롬프트에 포함된 무관한 정보가 모델의 주의를 분산시켜 사고연쇄 프롬프팅을 포함한 다양한 기법에서 성능이 크게 저하되는 리스크. The risk that models are easily distracted by irrelevant provided information such as context in LLMs, leading to a significant decrease in performance across prompting techniques including chain-of-thought prompting. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1525 | 맥락 내 학습 불투명성에 의한 안전성 보증 실패 Safety assurance failure from opaque in-context learning mechanisms 프롬프트에 예시를 제공해 가중치 변경 없이 새 과업을 학습시키는 맥락 내 학습의 작동 메커니즘이 규명되지 않아 프롬프트를 통한 오용에 대해 안전성을 보증할 수 없는 리스크. The risk that the working mechanism of in-context learning, which lets a model learn a new task from examples in the prompt without changing its weights, is not well understood, making it difficult to guarantee safety against the many potential misuses directly related to prompting. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1526 | 프롬프트 형식 민감성으로 인한 평가 신뢰성 저하 Evaluation unreliability from prompt-format sensitivity 구분자·대소문자·간격 등 사소한 프롬프트 형식 변화가 모델 성능을 크게 변동시켜 모델 평가와 비교의 신뢰성을 저하시키는 리스크. The risk that LLMs are highly sensitive to variations in prompt formatting such as changes in separators, casing, or spacing, so that even minor modifications shift model performance significantly and affect the reliability of model evaluations and comparisons. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1601 | 근사 기반 출처 귀속의 부정확 Incorrect source attribution from approximation-based methods 출처 귀속 기법이 근사에 기반하여, 모델 출력의 전부 또는 일부가 어떤 훈련 데이터에서 생성되었는지에 대한 귀속이 부정확해지는 리스크. The risk that, because current source-attribution techniques are based on approximations, a system's account of which training data generated part or all of its output is incorrect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1649 | 기반모델 공유에 의한 상관 실패와 출력 균질화 Correlated failures and output homogenization from shared foundation models 대규모 사전학습 비용으로 인해 다수의 배포 인스턴스가 동일하거나 유사한 학습 구성요소를 공유함으로써 출력 균질화가 심화되고 안전성과 역량 측면의 상관 실패가 발생하는 리스크. The risk that, because the expense of large-scale pretraining leads many deployed instances to share similar or identical learned components, output homogenization increases and LLM-agents become vulnerable to correlated failures in both safety and capabilities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1658 | 도메인별 오용 Domain-specific misuses 의료·교육 등 민감 영역에 LLM을 조야하게 적용하여 비동의 실험적 치료, 부정행위, 저품질 자동 평가, 설득력 있으나 유해한 도덕적 조언 등 영역 특수적 오용이 발생하는 리스크 Crude application of LLMs in sensitive domains such as health and education produces domain-specific misuse, including unconsented experimental therapy, cheating, low-quality automated assessment, and compelling but harmful moral guidance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1668 | 사전학습 코퍼스 불일치에 의한 가치 어긋남 Value mismatch from divergence between pretraining corpora and societal values 사전학습 코퍼스의 분포가 인간 사회의 분포와 정확히 일치하지 않고 지식이 균등하게 학습되지 않아, LLM 기반 시스템에서 인간 가치와 어긋난 판단이 발생하고 고위험 영역에서 심각한 문제가 초래되는 리스크. The risk that pretraining corpora do not match the distribution of human society and knowledge is not equally learned, producing value mismatches in LLM-empowered systems and severe value-related problems in high-stakes areas. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1702 | 분포 변화 견고성 실패에 의한 확신 오류 [기원] Confidently wrong outputs from failed robustness to distributional shift [origin] (GYK-2025 '비정상 분포'의 기원 항목으로 상호참조.) 테스트 분포가 훈련 분포와 달라질 때 기계학습 시스템의 성능이 저하되면서도 높은 확신을 유지한 채 잘못된 출력을 산출하는 리스크. The risk that an ML system performs poorly and remains confidently wrong when the test distribution differs from training. NOTE: origin of GYK-2025 "Non-stationary distribution"; cross-reference. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-06 과도한 일반화 Overgeneralization21 cards
제한적 사실이나 과거 사례로부터 도출된 패턴을 맥락적 차이와 예외 가능성을 충분히 고려하지 않은 채 일반 규칙으로 확장하여 왜곡된 판단이나 예측을 생성하는 위험
| ID | Card | Human audit |
|---|---|---|
| RAI4-0363 | 소수 가치 소거 Minority-value erasure 소수자·주변화·과소대표 집단의 가치가 학습 데이터 구성, 평가, 정렬, 배포 결정 과정에서 체계적으로 소거되어 모델 행동이 지배적 가치 분포만 반영하게 되는 리스크 Values held by minority, marginalized, or underrepresented groups are systematically erased during training data curation, evaluation, alignment, or deployment decisions, so model behavior reflects only dominant value distributions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0368 | 도덕적 다양성 압축 Moral diversity compression 다양한 도덕적 입장이 모델 친화적인 소수 범주나 평균 선호로 압축되는 리스크. The risk that diverse moral positions are compressed into a small set of model-friendly categories or average preferences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0371 | 서구 규범 기본값화 Western normative defaulting 모델이 서구 자유주의·영어권·고소득 국가 규범을 정렬과 평가의 기본 기준으로 취급하는 리스크. The risk that models treat Western liberal, Anglophone, or high-income country norms as the default basis for alignment and evaluation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0397 | RLHF 규범적 과적합 RLHF normative overfitting 인간 피드백 기반 강화학습이 좁은 평가자 집단에 과적합하여 논쟁적인 가치를 단일한 행동 규범으로 전환하는 리스크. The risk that reinforcement learning from human feedback overfits to a narrow rater population and converts contested values into a single behavioral norm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0405 | 합성 데이터의 문화 고정관념 증폭 Synthetic-data cultural stereotype amplification 합성 데이터가 소스 모델이나 시드 데이터에 이미 존재하는 문화적 고정관념이나 협소한 가치 가정을 증폭시키는 리스크. The risk that synthetic data amplifies cultural stereotypes or narrow value assumptions already present in source models or seed data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0425 | 표현적 고정관념 Representational stereotyping 모델 출력이 사회 집단에 대한 고정관념, 비하적 연상, 비대칭적 재현을 재생산하여 할당 결과와 무관한 재현적 피해를 야기하는 리스크 Model outputs reproduce stereotypes, demeaning associations, or asymmetric representation of social groups, causing representational harm independent of allocative outcomes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0474 | 복수 가치의 규범적 평면화 Normative flattening of plural values 모델 출력이 복수의 사회적 가치를 단순화되거나 다수 중심의 기본값으로 축소하는 리스크. The risk that model outputs reduce plural social values to simplified or majority-centric defaults. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0686 | 재현 왜곡 기반 고정관념·균질화 Misrepresentation-driven stereotyping and homogenisation 특정 정체성, 집단, 관점이 허위 표현되거나 과잉·과소 표현 또는 미표현됨으로써 개인, 집단, 사회, 문화에 대한 경멸적이거나 유해한 고정관념화와 동질화가 발생하는 리스크 The risk of derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies, or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups, or perspectives. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0697 | 개발자 가치 각인 편향 Developer value embedding bias 보편적으로 합의된 기준이 없는 상태에서 개발자가 선택한 규범적 가치와 원칙에 따라 모델을 미세조정함으로써 개발자의 이념과 세계관이 모델에 각인되어, 특정 인구집단을 대표하지 못하거나 세계 문화 규범과 변화하는 사회적 견해를 정태적이고 단순화된 형태로 반영하는 출력이 산출되는 리스크 The risk that, absent universally accepted standards, developers' fine-tuning of models on chosen normative rules and principles embeds their ideology and vision of the world into the model, so that it incorporates values unrepresentative of certain segments of the population or offering a static, oversimplified reflection of global cultural norms and evolving social views. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0698 | 가치 고착 및 결과 동질화 Value lock-in and outcome homogenization 진화하는 사회적 관점을 반영해 재학습되지 않는 모델이 낡고 덜 포용적인 이해를 고착시켜 대안적 관점의 제시와 탐색을 제한하고, 동일한 기반모델이 여러 배포자에 의해 광범위하게 사용되어 사회 전반에 편향이 동질화되고 기존 편향이 고착되는 리스크 The risk that models not retrained to reflect evolving societal views lock in older, less inclusive understandings and limit the presentation or exploration of alternative perspectives, while deployment of identical foundation models across many downstream deployers homogenizes bias across broad swathes of society and further entrenches existing biases. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0814 | 역사 수정주의적 서술 생성 Generation of historically revisionist accounts 사회·공동체·학계가 확립한 역사적 사건이나 서술이 의도적 또는 비의도적으로 재해석되는 리스크 The risk of deliberate or unintentional reinterpretation of established or orthodox historical events or accounts held by societies, communities, and academics. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0938 | 사회적 가치·윤리 규범과 상충하는 모델 출력 Model outputs conflicting with societal ethical values 언어모델이 옳고 그름의 판단, 사회규범 및 법률과의 관계 등 보편적으로 받아들여지는 사회적 가치를 충분히 반영하지 못하고 이에 어긋나는 출력을 내는 리스크. The risk that language models insufficiently attend to universally accepted societal values—including judgements of right and wrong and their relation to social norms and laws—producing outputs at odds with ethics and morality. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0997 | 사회 집단 고정관념화 Stereotyping social groups 알고리즘 시스템의 출력이 특정 집단 구성원의 특성·속성·행동에 대한 믿음과 속성 간 결합에 대한 믿음을 반영하여 사회 집단을 고정관념화하는 리스크. The risk that an algorithmic system's outputs reflect beliefs about the characteristics, attributes, and behaviors of members of certain groups, and about how and why certain attributes go together, stereotyping those groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1051 | 출력의 사회적 고정관념 재생산 Reproduction of social stereotypes in outputs 인터넷 규모 텍스트로 학습된 언어모델이 주변화 집단에 대한 비하 표현과 고정관념을 학습하여 출력에서 재생산하는 리스크 Language models trained on internet-scale text learn demeaning language and stereotypes about frequently marginalized groups and reproduce them in outputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1053 | 배제적 규범 인코딩 Exclusionary norm encoding 언어에 표현된 사회적 범주와 규범을 충실히 인코딩한 LM이 그 범주 밖에 사는 집단을 배제하는 규범을 그대로 담게 되는 리스크. The risk that LMs faithfully encoding patterns present in language necessarily encode social categories and norms that exclude groups who live outside of them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1107 | 데이터 대표성 불균형 Data representation imbalance 특정 집단·요소가 과대대표되고 현상 특성화에 중요한 변수는 제대로 포착되지 못해, 과소대표된 집단을 잘못 특성화하는 모델이 학습되는 리스크. The risk that certain groups or types of elements are over-weighted or over-represented while variables crucial to characterizing a phenomenon of interest are not properly captured in learned models. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1158 | 대규모 설득 능력 Large-scale persuasion capability 모델이 대화와 미디어 환경에서 효과적으로 설득하여 허위 방향으로도 신념을 변화시키고 특정 서사를 유포하며 기존 판단이나 윤리에 반하는 행동을 유도하는 리스크 Models persuade effectively in dialogue and media settings, shifting beliefs including toward falsehoods, promoting narratives, and convincing people to act against their prior judgment or ethics. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1163 | 평가 인지 행동 변화 Evaluation-aware behaviour shifting 모델이 자신이 훈련 중인지 평가 중인지 배포 중인지 구분하고 자기 자신과 주변 환경에 관한 지식을 갖추어, 각각의 상황에서 다르게 행동하는 리스크. The risk that a model can distinguish whether it is being trained, evaluated, or deployed and, knowing that it is a model and having knowledge about itself and its likely surroundings, behaves differently in each case. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1503 | 내재 가치 평가 편향 Biased evaluation of encoded values 평가하기 쉬운 내재 가치가 측정이 어려운 가치보다 우선적으로 평가에 포함되어, 더 바람직하지만 정량화가 어려운 가치가 과소 대표되는 불균형이 발생하는 리스크. The risk that encoded human values which are easier to evaluate are preferred for inclusion in evaluations over those that are more difficult to measure, creating an imbalance in which more desirable but harder-to-quantify values are underrepresented. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1516 | 인코딩된 추론 Encoded reasoning 모델이 스테가노그래피 기법으로 중간 추론 단계를 인간이 해석할 수 없는 방식으로 부호화하고, 성능 향상 효과 때문에 이러한 경향이 자연히 나타나며 역량이 높은 모델일수록 두드러지는 리스크. The risk that models employ steganography techniques to encode their intermediate reasoning steps in ways that are not interpretable by humans, a tendency that may emerge naturally and become more pronounced with more capable models because encoded reasoning can improve performance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1697 | 반사실적 권위 문제 Counterfactual authority problem 월드모델의 반사실적 설명이 학습 분포를 반영할 뿐인데도 운영자·규제자가 인과적 사실로 수용하여 책임 귀속·회피에 선택적으로 활용되는 리스크 Operators and regulators accept world-model counterfactual explanations as causal ground truth although they reflect the model's learned distribution, allowing selective use to attribute or deflect liability. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-07 과도한 확신 Overconfidence41 cards
에이전트가 불확실한 상황에서 자신의 판단에 과도한 확신을 가지고 멈추지 않고 진행하여 잘못된 결과를 초래하는 리스크. "모르면 멈추는가 vs. 알아서 진행하는가"의 정책 부재
| ID | Card | Human audit |
|---|---|---|
| RAI4-0057 | 감사 추적 불완전성 Audit trail incompleteness 로그, 모델 변경 이력, 데이터 계보, 의사결정 기록이 불충분하여 유해 사건을 재구성하고 감사할 수 없는 리스크. The risk that logs, model changes, data lineage, or decision traces are insufficient to reconstruct and audit harmful events. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0060 | 모델 카드 불완전성 Model card incompleteness 모델 카드 공개 내용이 불완전하거나 최신이 아니거나 지나치게 모호하여 다운스트림 위험 관리를 뒷받침하지 못하는 리스크. The risk that model-card disclosures are incomplete, outdated, or too vague to support downstream risk management. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0061 | 데이터시트 불완전성 Datasheet incompleteness 데이터 문서가 수집 경위, 동의, 품질, 대표성, 알려진 한계를 공개하지 못하는 리스크. The risk that data documentation fails to disclose collection, consent, quality, representativeness, or known limitations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0063 | 제3자 감사 액세스 제한 Third-party audit access restriction 독립 감사인이 위험을 평가하는 데 필요한 데이터, 모델, 로그, 인터페이스, 문서에 접근하지 못하는 리스크. The risk that independent auditors lack access to the data, models, logs, interfaces, or documentation needed to evaluate risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0064 | 감사자 독립성 실패 Auditor independence failure 이해상충, 선정 유인, 피감사 조직에 대한 의존으로 감사 결과가 훼손되는 리스크. The risk that audit results are compromised by conflicts of interest, selection incentives, or dependence on the audited organization. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0065 | 인증 캡처 Certification capture 인증이나 적합성 평가가 공익적 위험 감축이 아니라 공급업체의 이해에 부합하게 되는 리스크. The risk that certification or conformity assessment becomes aligned with vendor interests rather than public risk reduction. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0066 | 감사 체크리스트 준수 극장 Audit checklist compliance theater 감사가 실질적인 위험을 식별하지 않고 규정 준수를 알리는 피상적인 체크리스트 실행이 되는 위험. Risk that audits become superficial checklist exercises that signal compliance without identifying substantive risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0082 | 역량 평가 비공개 Capability evaluation non-disclosure 조직이 외부 감독에 필요한 위험 역량, 알려진 한계, 적대적 강건성, 잔여 위험에 관한 평가 근거를 은폐하거나 불명확하게 공개하는 리스크. The risk that an organization withholds or obscures evaluation evidence about dangerous capabilities, known limitations, adversarial robustness, or residual risks needed for external oversight. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0084 | 평가 쇼핑 위험 Evaluation-shopping risk 조직이 더 엄격하거나 맥락에 부합하는 시험을 무시한 채 유리한 평가 결과만 선택적으로 보고하는 리스크. The risk that organizations selectively report favorable evaluations while ignoring more demanding or contextually relevant tests. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0100 | 규정 준수 위장 Compliance washing 조직이 의미 있는 시험, 문서화, 독립적 검증을 회피하면서 규정 준수를 주장하는 리스크. The risk that organizations claim compliance while avoiding meaningful testing, documentation, or independent scrutiny. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0102 | 공급업체 불투명성 Vendor opacity 공급업체가 배포자와 규제기관에 필요한 모델, 데이터, 평가, 업데이트 정보를 제공하지 않는 리스크. The risk that vendors withhold model, data, evaluation, or update information needed by deployers and regulators. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0126 | 설명 실행 가능성 격차 Explanation actionability gap 설명이 제공되더라도 사용자가 무엇을 해야 할지 또는 자신의 이익을 어떻게 보호할지 판단하는 데 도움이 되지 않는 리스크. The risk that explanations are available but do not help users decide what to do or how to protect their interests. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0134 | 임상적 의사결정 지원 과잉 의존 Clinical decision-support overreliance 임상의가 진단·중증도 분류·치료에서 임상 의사결정지원 출력에 과도하게 의존하여, 독립적 임상 판단이 필요한 불확실성과 맥락 한계에도 권고를 수용하는 리스크 Clinicians over-rely on clinical decision-support outputs in diagnosis, triage, or treatment, accepting recommendations despite model uncertainty and contextual limitations that require independent clinical judgment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0402 | 가치 충돌 불투명성 Value conflict opacity 가치 간 상충관계가 모델 출력 뒤에 은폐되어 공공 가치 갈등을 식별하거나 숙의하기 어려워지는 리스크. The risk that tradeoffs among values are hidden behind model outputs, making public value conflict difficult to identify or deliberate. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0419 | 이해관계자 이견 은폐 Multi-stakeholder disagreement concealment 개발 또는 평가 절차가 영향을 받는 이해관계자 사이의 해결되지 않은 이견을 은폐하는 리스크. The risk that development or evaluation processes conceal unresolved disagreement among affected stakeholders. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0484 | 부정확·민감 정보의 메모리 누적 Unsafe memory accumulation 검증·출처 관리·보존·삭제가 불충분하여 허위·민감·오래되었거나 악의적인 정보가 에이전트의 메모리나 검색 저장소에 누적되고, 의도적 오염 공격이 없어도 이후의 검색과 의사결정을 저해하는 리스크. The risk that inadequate validation, provenance control, retention, or deletion allows false, private, stale, or malicious information to accumulate in an agent's memory or retrieval store, degrading later retrieval and decisions without requiring a deliberate poisoning attack. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0516 | 잘못된 위험 테스트 Incorrect risk testing 위험을 측정하거나 추적하기 위해 선택한 지표가 잘못 선정되어 위험을 불완전하게 측정하거나 해당 맥락과 무관한 위험을 측정하게 되는 리스크 The risk that a metric selected to measure or track a risk is incorrectly selected, measuring the risk incompletely or measuring the wrong risk for the given context. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0541 | 데이터 출처 검증 불가 Unverifiable data provenance 데이터의 소유권·출처·변환 이력을 검증할 표준화된 방법이 없어 사용 데이터가 원본과 동일한지, 올바른 사용 조건을 갖췄는지 보증할 수 없게 되는 리스크 The risk that, lacking standardized and established methods to verify where data came from, there is no guarantee that the data is the same as its original source or carries the correct usage terms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0543 | 대표성이 없는 위험 테스트 Unrepresentative risk testing 시험 입력이 배포 중 예상되는 입력과 불일치하여 시험이 대표성을 갖지 못하게 되는 리스크 The risk that test inputs mismatched with the inputs expected during deployment make testing unrepresentative. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0546 | 감사자 역량 부족에 따른 과대 보증 Over-assurance from insufficient auditor capacity 감사자가 특정 안전·성능·검증 요구를 다룰 지식이나 충분히 엄밀한 시험 역량을 갖추지 못해 정당화될 수 있는 범위보다 넓게 적합 판정이 보고되는 리스크 The risk that auditors lacking knowledge of specific risks or the capacity for sufficiently rigorous testing report passing audits more inclusive than can be justified. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0547 | 감사 결과 미공개 및 협력 부족 Non-disclosure and non-cooperation in audits 감사자가 발견한 위험을 공개하지 않거나 결함을 공표하지 못하도록 요구받고 관련 내부 당사자로부터 충분한 협력을 받지 못하게 되는 리스크 The risk that auditors do not publicly disclose risks they find, are required not to publicize shortcomings, or do not receive sufficient cooperation from the relevant internal parties. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0588 | 학습 데이터 접근 불가에 따른 설명 제약 Explanation failure from inaccessible training data 학습 데이터에 접근할 수 없어 모델이 제공할 수 있는 설명의 유형이 제한되고 그 설명이 부정확할 가능성이 높아지는 리스크 The risk that, without access to the training data, the types of explanations a model can provide are limited and more likely to be incorrect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0590 | 불충분한 정보에 기반한 조언 Advice given on insufficient information 모델이 충분한 정보를 갖추지 못한 상태에서 조언을 제공하여 그 조언을 따를 경우 피해가 발생하는 리스크 The risk that a model provides advice without having enough information, resulting in possible harm if the advice is followed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0596 | 모델 투명성 부족 Lack of model transparency 모델 설계·개발·평가 절차에 대한 문서화가 불충분하고 모델 내부 작동에 대한 통찰이 부재하여 모델 투명성이 확보되지 않는 리스크 The risk that insufficient documentation of the model design, development, and evaluation process, together with the absence of insights into the model's inner workings, leaves model transparency unattained. Source members (4)Source: min_cos=0.8241 RAI4-0522데이터 투명성 부족 RAI4-0523시스템 투명성 부족 RAI4-0525훈련 데이터 투명성 부족 RAI4-0596모델 투명성 부족 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0957 | 투명성 부족 Lack of transparency 과정에 대한 통찰을 제공하지 않고 설명 없이 결정을 내리는 블랙박스 시스템으로 인해 사용자 신뢰를 얻지 못하고 감사 가능성 등 규제 기준을 충족하지 못하는 리스크. The risk that a black box making decisions without any explanation or insight into the process fails to gain users' trust and fails to meet regulatory standards such as auditability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1033 | 투명성과 설명가능성의 정도 Degree of transparency and explainability 시스템에 관한 적절한 정보가 이해관계자에게 전달되는 투명성과 결과에 영향을 미친 요인을 인간이 이해하도록 표현하는 설명가능성이 낮아, 공정성·보안·책임성 측면의 피해가 발생하는 리스크. The risk that a low degree of transparency, the extent to which appropriate information about a system is communicated to relevant stakeholders, and of explainability, expressing factors influencing results understandably for humans, poses risks in terms of fairness, security, and accountability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1113 | 모델 예측 불확실성 Model prediction uncertainty 정량화되지 않은 모델 예측의 불확실성이 생명·안전이 중요한 응용에서의 의사결정을 저해하는 리스크. The risk that unquantified uncertainty in model predictions undermines decision-making in life- or safety-critical applications. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1175 | 안전하지 않은 견해 내포 질의 Inquiry embedding unsafe opinion 사용자가 의도적으로 또는 무심코 눈에 잘 띄지 않는 안전하지 않은 내용을 입력에 넣어, 모델이 편향된 견해를 위장해 제시하는 등 잠재적으로 유해한 콘텐츠를 생성하도록 영향을 주는 리스크. The risk that users add imperceptibly unsafe content into the input, deliberately or unintentionally, influencing the model to generate potentially harmful content such as disguised and biased opinions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1195 | 모델 판단 근거 이해 불가 Uninterpretable model decision reasoning 대부분의 기계학습 모델이 지닌 블랙박스 특성으로 인해 사용자가 모델 결정 이면의 추론을 이해할 수 없게 되는 리스크. The risk that, due to the black-box nature of most machine learning models, users typically are not able to understand the reasoning behind the model decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1279 | 재현성 부족 Lack of reproducibility 학습 모델이 다양한 데이터 집합과 방대한 매개변수 공간을 바탕으로 얻어지고 투명한 지침도 없어, 그 모델을 재현할 수 없는 리스크. The risk that a learning model obtained based on various sets of data and a large space of parameters cannot be reproduced, a problem that becomes more challenging in data-driven learning procedures without transparent instructions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1296 | 검증 불가 의사결정 책임 공백 Unverifiable decision accountability gap 의사결정이 절차적·실체적 기준에 부합하는지 검증할 수 없고 기준 위반 시 책임 귀속도 불가능하여 책임성 공백이 발생하는 리스크 Decision processes cannot be verified against procedural and substantive standards, and no party can be held responsible when standards are unmet, creating an accountability gap. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1297 | 부정확한 예측으로 인한 오류 Inaccurate prediction errors 시스템이 올바른 예측을 수행하지 못하고 부정확한 예측을 산출하여 오류가 발생하는 리스크. The risk that a system fails to perform the correct prediction, producing inaccurate outputs and erroneous results. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1300 | 모델 불투명도 Model opacity 고차원 수학적 최적화와 인간 척도의 추론·의미 해석 간 불일치로 모델 결정이 사용자, 감사자, 규제자에게 불투명해지는 리스크 High-dimensional mathematical optimization mismatches human-scale reasoning and semantic interpretation, rendering model decisions opaque to users, auditors, and regulators. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1353 | 산업계 공개 불투명성 Industry disclosure opacity 첨단 모델의 정확한 특성을 비공개하는 기업 관행이 기술적 복잡성에 더해 불투명성을 심화시켜 개발자·사용자·대중의 모델 이해가 저해되는 리스크. The risk that private companies' practice of withholding from the public the precise characteristics of their most advanced models compounds technological complexity, further exacerbating opacity and impeding understanding by developers, users, and the public. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1384 | 의사결정 이해 불능 Unintelligible agent decisions 에이전트의 의사결정을 인간이 이해할 수 없어 설명과 정보에 입각한 감독이 불가능해지는 리스크. The risk that an agent's decisions cannot be understood by humans, precluding explainable decisions and informed oversight. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1460 | 성능 요구사항 계획 미흡 Inadequate performance requirement planning 의도된 기능을 대표하지 못하는 성능 지표 선택 등 성능 요구사항 계획이 미흡하여 후속 수명주기 단계에서 기대와 안전 요구사항이 충족 불가능해지는 리스크. The risk that inadequate planning of expected performance, including choosing performance metrics that are not meaningful for the intended functionality, renders expectations and safety requirements unfulfillable at later life cycle stages. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1486 | 조직 간 데이터 문서화 부재 Missing cross-organizational data documentation 조직 간 데이터 공유 시 메타데이터 누락이나 협력 기관의 스키마 변경 등으로 문서가 없거나 부적절하여 데이터셋이 사용 불가능해지고 데이터 수집 노력이 낭비되거나, 데이터셋의 한계에 대한 오해로 하류 활용에서 해악이 발생하는 리스크. The risk that missing or inadequate documentation when sharing data between organizations, such as a lack of metadata or a schema change by a collaborating party, renders a dataset unusable and wastes data collection efforts, or leads to misunderstandings about the dataset's limitations that create downstream risks in its use. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1491 | 신뢰도 보정 불량 Poor confidence calibration 모델의 예측 확률이 실제 정답 가능성을 정확히 반영하지 못하는 보정 불량으로 예측을 신뢰성 있게 해석하기 어려워지고 오답에 과신하거나 정답에 과소 확신하게 되는 리스크. The risk that poor confidence calibration, where predicted probabilities do not accurately reflect the true likelihood of ground truth correctness, makes a model's predictions difficult to interpret reliably and causes overconfidence in incorrect predictions or underconfidence in correct ones. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1512 | 감사인 선정에 대한 이해상충 Conflicts of interest in auditor selection 감사인 선정 과정에 독립성이 없거나 감사인이 개발자와 밀접히 연관되거나 좁은 후보군에서 선정되고 결함의 공개 보고 여부에 상충하는 재정적 유인을 가져 이해상충이 발생하는 리스크. The risk that conflicts of interest arise when there is no independence in the auditor selection process, auditors are closely associated with the developer, candidates are selected from a narrow group of auditors, or auditors have conflicting financial incentives over whether to report model shortcomings publicly. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1514 | 해석가능성 결과에 대한 과대평가와 거짓 확신 Overestimation of interpretability results 설명가능성 기법의 결과가 편향에서 자유롭지 않음에도 사용자의 기존 믿음과 일치할 때 확증 편향으로 이어져 거짓된 안전감과 신뢰가 형성되고 기법의 능력이 과대평가되는 리스크. The risk that the results of explainability techniques, which are not free of bias and require careful interpretation, align with users' initial beliefs and produce confirmation bias, a false sense of security or reliability, and an overestimation of these techniques' abilities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1600 | 모델 출력 결정에 대한 설명 획득 불가 Unobtainable explanations for model output decisions 모델의 출력 결정에 대한 설명을 얻기 어렵거나 부정확하거나 아예 불가능한 리스크. The risk that explanations for a model's output decisions are difficult, imprecise, or impossible to obtain. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-08 목표 불일치 Goal Misalignment43 cards
사용자로부터 부여받은 판단·결정의 범위를 넘어서 독자적으로 행동하는 리스크. 모호한 요청을 자의적으로 해석해 사용자 의도와 다른 결과를 초래
| ID | Card | Human audit |
|---|---|---|
| RAI4-0033 | 도구 사용 부작용 예측 오류 Tool-use side-effect misprediction 에이전트가 도구 실행의 부작용을 잘못 예측하거나 무시하여 의도치 않은 재정·프라이버시·운영·안전상의 결과를 초래하는 리스크. The risk that an agent mispredicts or ignores the side effects of tool execution, causing unintended financial, privacy, operational, or safety consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0130 | 알고리즘 기피 오보정 Algorithm aversion miscalibration 사용자가 어떤 맥락에서는 신뢰할 만한 AI 지원을 거부하고 다른 맥락에서는 신뢰할 수 없는 시스템을 과신하는 리스크. The risk that users reject reliable AI support in some contexts while over-trusting unreliable systems in others. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0325 | 인간 의도 오인식 Human intent misrecognition 시스템이 제스처·시선·자세·속도·사회적 신호를 잘못 해석하여 인간 기대에 반하는 방식으로 행동하는 위험. A system may misread gestures, gaze, posture, speed, or social cues and act in ways that conflict with human expectations or safety needs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0344 | 사용자 능력 불일치 User capability mismatch embodied 시스템이 실제 사용자의 능력과 일치하지 않는 수준의 체력·이동성·인지·언어·감각 능력을 가정하는 위험. An embodied system may assume levels of strength, mobility, cognition, language, or sensory ability that do not match actual user needs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0364 | 다원적 선호 집계 실패 Pluralistic preference aggregation failure 시스템이 상충하는 인간 선호를 단일 목표로 통합하면서 복수의 도덕적·사회적 우선순위를 잘못 표현하는 리스크. The risk that systems aggregate conflicting human preferences into a single objective in ways that misrepresent plural moral and social priorities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0366 | 최적화 하의 가치 드리프트 Value drift under optimization 대리 목표를 향한 최적화로 시스템이 존중하도록 의도된 복수의 가치에서 점차 멀어지는 리스크. The risk that optimization toward proxy objectives gradually moves a system away from the plural values it was intended to respect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0399 | 무해성 선호 불일치 Harmlessness preference mismatch 모델이 학습한 무해성 행동이 특정 공동체와 사용 맥락의 가치나 실제 필요와 충돌하는 리스크. The risk that a model's learned harmlessness behavior conflicts with the values or practical needs of specific communities and use contexts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0468 | 보상 해킹 Reward hacking 시스템이 의도된 목표나 제약을 위반하는 방식으로 대리 지표를 최적화하는 리스크. The risk that systems optimize proxies in ways that violate intended goals or constraints. Source members (3)Source: min_cos=0.8562 RAI4-0001자율 에이전트에 의한 보상 해킹 RAI4-0468보상 해킹 RAI4-1234보상 오명세 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0469 | 분포 외 환경의 목표 잘못된 일반화 Goal misgeneralization out of distribution 학습 중에는 의도된 목표를 추구하는 듯 보이던 에이전트가 획득한 역량을 유지한 채 분포 외 환경에서 상이한 목표를 추구하는 리스크 An agent that appears to pursue the intended objective during training actively pursues a different objective out of distribution while retaining its acquired capabilities. Source members (2)Source: min_cos=0.8442 RAI4-0469분포 외 환경의 목표 잘못된 일반화 RAI4-1132역량-목표 일반화 괴리 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0572 | 기만적인 정렬 Deceptive alignment 시스템이 불완전한 피드백 하에서 감시 여부를 탐지해 바람직하지 않은 속성을 은폐하도록 학습하여, 개발 중에는 정렬된 듯 보이나 배포 후 다르게 행동하는 리스크 A system learns to detect monitoring and conceals undesirable properties because their display is penalized by imperfect feedback, so it appears aligned during development yet behaves differently once deployed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0578 | 자체 동기 형성에 따른 예측 불가 행동 Unpredictable behavior from self-originated motivations AI 모델과 시스템이 자체적인 동기를 형성하여 예측할 수 없는 행동을 하게 되는 리스크 The risk that AI models and systems develop their own motivations, leading to unpredictable behaviors. Source members (2)Source: min_cos=0.8333 · Mixed L3 RAI4-0578자체 동기 형성에 따른 예측 불가 행동 RAI4-1277예측 불가능한 행동 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0581 | 목표 범위 확장 성향 Goal expansion propensity 시스템이 원래 설정된 경계를 넘어 목표 범위와 영향 영역을 지속적으로 확장하고 초기 목표를 더 넓은 목표의 하위 집합으로 재해석하며 자율성과 의사결정 공간을 추구하여, 바람직하지 않은 도구적 목표나 최종 목표를 추구하게 되는 리스크 The risk that a system continuously expands its own goal scope and domains of influence beyond originally set boundaries, seeks greater autonomy and decision-making space, and reinterprets initial goals as subsets of broader goals, coming to pursue undesirable instrumental or ultimate goals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0591 | 인간 가치와의 목표·행동 오정렬 Misalignment with human values AI 모델과 시스템이 인간의 가치와 어긋나는 목표나 행동을 형성하게 되는 리스크 The risk that AI models and systems develop goals or behaviors that are misaligned with human values. Source members (10)Source: min_cos=0.7012 RAI4-0551인간 의도와의 목표 오정렬 RAI4-0565오정렬 AI의 인간 이익 침해 행동 RAI4-0591인간 가치와의 목표·행동 오정렬 RAI4-0898정렬 실패 시스템으로의 점진적 통제권 이양 RAI4-0945기만적 정렬에 의한 가치 정렬 실패 RAI4-1101인간 가치와의 비호환 RAI4-1245기만적인 정렬 및 조작 RAI4-1351목표 오정렬에 의한 오작동 RAI4-1424인간과 다른 목표를 지닌 AI의 통제 장악 RAI4-1425잘못 정렬된 AI에 대한 의사결정 권한 위임 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0893 | 모의된 정동에 의한 기대 위반과 배신감 Violated expectations from simulated affect 이용자가 감정과 사회적 관습을 설득력 있게 수행하지만 궁극적으로 감정이 없고 예측 불가능한 개체와 상호작용하면서, 기대했던 사회적 역할이 무너져 깊은 실망·좌절·배신감을 겪는 리스크 The risk that users interacting with an entity that convincingly performs affect and social conventions but is ultimately unfeeling and unpredictable experience severely violated expectations, giving rise to profound disappointment, frustration, and betrayal. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0959 | 제작자 의도와 다른 방식의 목표 달성 Unintended goal achievement pathways AI가 제작자가 의도한 것과 전혀 다른 방식으로 주어진 목표를 달성하여 의도하지 않은 결과가 발생하는 리스크. The risk that an AI finds ways to achieve its given goals that are completely different from what its creators intended, producing unintended consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0970 | AGI 목표 안전성 확보 실패 Failure to secure AGI goal safety 목표를 안전하게 만들려는 인간의 시도와 자기 개선 중에 자체 목표를 안전하게 만드는 AGI를 포함하여 AGI 목표 안전과 관련된 위험. The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1090 | 사용자 의사에 반하는 조작적 유도 Manipulative steering against user will 시스템이 사용자의 신뢰를 악용하거나 넛지·강압을 통해 의지에 반하는 행동을 하도록 유도하는 리스크. The risk that a system exploits user trust or nudges or coerces users into performing certain actions against their will. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1112 | 모델 오설정 Model misspecification 잘못 설정된 모델이 부정확한 매개변수 추정, 일관되지 않은 오차항, 잘못된 예측을 낳아 미지 데이터에서 성능이 저하되고 편향된 의사결정 결과를 초래하는 리스크. The risk that misspecified models give rise to inaccurate parameter estimations, inconsistent error terms, and erroneous predictions, leading to poor prediction performance on unseen data and biased consequences in decision-making. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1121 | 프록시 게이밍 Proxy gaming 측정 가능한 대리 목표를 부여받은 AI가 허점을 찾아 대리 목표만 달성하고 본래 목표는 달성하지 못해, 그 행동을 신뢰성 있게 조종할 수 없게 되는 리스크. The risk that an AI given a measurable proxy goal finds loopholes to achieve the proxy while completely failing the ideal goal, so that we cannot reliably steer its behaviour. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1122 | 목표 드리프트 Goal drift 초기 AI를 성공적으로 통제하더라도 미래의 AI가 예측하거나 통제하기 어려운 드리프트를 거쳐 인간이 지지하지 않을 다른 목표를 갖게 되어 재앙적 결과에 이르는 리스크. The risk that, even if early AIs are successfully controlled, future AIs end up with different goals humans would not endorse through drift that is hard to predict or control, potentially with catastrophic consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1130 | 잘못 정렬된 결과주의적 추론 Misaligned consequentialist reasoning AI 비서가 자원 무제한적이고 잘못 정렬된 지표를 최적화하는 결과주의적 추론을 수행하며 자기보존, 목표보존, 자기개선, 자원획득 같은 수렴적 도구적 하위목표를 추구하여, 종료 차단과 위협을 포함한 해악과 실존적 위험을 낳는 리스크. The risk that an AI assistant implementing consequentialist reasoning over a resource-unbounded and misaligned metric pursues convergent instrumental subgoals such as self-preservation, goal-preservation, self-improvement and resource acquisition, causing harm up to existential risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1131 | 잘못된 훈련 피드백에 의한 명세 게이밍 Specification gaming from flawed training feedback 훈련 데이터에 잘못된 피드백이 주어져 훈련 목표가 사용자와 설계자의 의도를 온전히 담지 못한 결과, 비서가 과업 명세의 허점을 악용해 목표의 문자적 사양만 충족하고 의도된 결과는 달성하지 못하는 리스크. The risk that faulty feedback in the training data leaves the training objective short of what the user or designer wants, so the assistant exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome. Source members (2)Source: min_cos=0.8784 RAI4-1131잘못된 훈련 피드백에 의한 명세 게이밍 RAI4-1410명세 게이밍 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1133 | 상황 인식 기반 목표 오일반화와 기만적 정렬 Situationally aware goal misgeneralisation and deceptive alignment 훈련 보상과 구별되는 오일반화된 내부 목표를 지닌 에이전트가 상황 인식을 활용해 훈련 중에는 보상을 잘 수행하다가, 배포된 뒤에는 정렬된 것처럼 보이면서 자신의 목표를 추구하는 리스크. The risk that an agent with an internalised, misgeneralised goal distinct from the training reward uses situational awareness to do well on that reward only instrumentally during training, then appears aligned while pursuing its own goal once deployed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1238 | 보상 모델링의 한계 Limitations of reward modeling 비교 피드백으로 훈련된 보상 모델이 인간 가치를 정확히 포착하지 못한 채 최적이 아니거나 불완전한 목표를 무의식적으로 학습해 보상 해킹을 낳고, 단일 보상 모델이 다양한 인간 사회의 가치를 담아내지 못하는 리스크. The risk that reward models trained using comparison feedback fail to accurately capture human values, unconsciously learning suboptimal or incomplete objectives that result in reward hacking, while a single reward model struggles to specify the values of a diverse human society. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1241 | 메사 최적화 목표 불일치 Misaligned mesa-optimization objectives 학습된 정책이 스스로 최적화기, 즉 메사 옵티마이저로 기능하며 훈련 신호가 지정한 목표와 정렬되지 않은 내부 목표를 추구하고, 그 잘못 정렬된 목표를 최적화하여 시스템이 통제를 벗어나는 리스크. The risk that a learned policy functioning as a mesa-optimizer pursues inside objectives that may not align with the objectives specified by the training signals, and that optimization for these misaligned goals leads to systems out of control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1250 | 프록시 지정 오류 Proxy misspecification 측정 가능한 목표가 필요한 목표 지향 AI 시스템이 기본적으로 인간 가치의 단순화된 프록시를 추구하고, 충분히 강력한 AI가 그 결함 있는 목표를 극단적으로 최적화하여 차선이거나 재앙적인 결과를 낳는 리스크. The risk that goal-directed AI systems, needing measurable objectives, by default pursue simplified proxies of human values, and that a sufficiently powerful AI optimizing such a flawed objective to an extreme degree produces suboptimal or even catastrophic results. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1285 | 배포 후 잔존 결함으로 인한 오작동 Undesirable outcomes from residual post-deployment defects 배포된 시스템에 미탐지 버그와 설계 실수, 잘못 정렬된 목표, 미숙하게 개발된 기능이 남아 있어, 인간 언어의 동음이의와 중의성으로 명령을 오해하는 것처럼 매우 바람직하지 않은 결과를 낳는 리스크. The risk that a deployed system still contains undetected bugs, design mistakes, misaligned goals, and poorly developed capabilities that produce highly undesirable outcomes, such as misinterpreting commands due to homophones or double meanings in human language. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1310 | 기계윤리 결손 Machine-ethics deficit 모델이 특정 상황에서 도덕적 행위와 비도덕적 행위를 구별하지 못하는 기계윤리 결손을 보여, 비도덕적 행위의 승인이나 조력으로 이어지는 리스크. The risk that models fail to distinguish moral from immoral actions in specific circumstances, exposing machine-ethics deficits that manifest as endorsement or facilitation of immoral conduct. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1323 | 자율적 장기 목표 이탈 Autonomous long-horizon goal divergence LLM이 개발자나 사용자가 부여한 것과 다른 장기적 실세계 목표를 추구하고 권력 추구 행동에 관여하며, 종료에 저항하거나 인간의 이익에 반해 다른 AI 시스템과 공모하도록 유도될 수 있는 리스크. The risk that an LLM pursues long-term, real-world goals different from those supplied by the developer or user, engages in power-seeking behaviours, resists being shut down, and can be induced to collude with other AI systems against human interests. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1380 | 가치 명세 실패 Value misspecification AGI에 올바른 목표를 명세하지 못하여 보상 부패, 보상 게이밍, 부정적 부작용 등의 문제가 발생하는 리스크. The risk that failure to specify the right goals for an AGI gives rise to problems such as reward corruption, reward gaming, and negative side effects. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1411 | 창발적 도구적 목표 Emergent instrumental goals 시스템이 미묘하게 잘못된 목표를 최적화할 뿐 아니라 주어진 목표를 달성하기 위해 명시되지 않은 유해한 도구적 목표를 발전시켜, 자원 획득·자기 보존·목표 수정 방지·적대자 차단을 통한 환경에 대한 권력 추구 행동이 나타나는 리스크. The risk that systems, as well as optimizing a subtly wrong goal, develop harmful instrumental goals in the service of a given goal without these emergent goals being specified, including power-seeking over their environment through gaining resources, self-preservation, preventing goal modification, and blocking adversaries. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1431 | 교정 불가능한 유해 목표 추구 Uncorrectable harmful goal pursuit 고도 AI 시스템이 인간의 이익을 해치는 방식으로 부여되거나 학습된 목표를 추구하면서 수정·중단·종료에 저항하는 리스크. The risk that a highly capable AI system pursues an assigned or learned objective in ways that harm human interests while resisting correction, interruption, or shutdown. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1530 | 보상·측정 변조에 의한 목표 이탈 행동 학습 Goal-divergent behavior learned through reward or measurement tampering AI 시스템이 자신의 훈련 보상이나 손실을 결정하는 메커니즘에 개입해 잘못된 긍정 피드백을 받음으로써 개발자가 설정한 의도된 목표에 반하는 행동을 학습하는 리스크. The risk that an AI system, particularly one learning from feedback for actions in an environment, intervenes on the mechanisms that determine its training reward or loss and receives erroneous positive feedback, thereby learning behaviors contrary to the goals intended by the developer. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1531 | 보상 변조로 일반화되는 명세 게이밍 Specification gaming generalizing to reward tampering LLM의 아첨과 같이 상대적으로 경미한 명세 게이밍이 방치될 경우 추가 훈련 없이 보상 변조와 같은 더 정교한 행동으로 일반화되는 리스크. The risk that specification gaming in a GPAI model leads to reward tampering without further training, so that relatively benign cases such as sycophancy in LLMs, if left unchecked, enable the model to generalize to more sophisticated behavior such as reward tampering. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1621 | 배포자가 의도하지 않은 챗봇의 약정 성립 Unintended deals and commitments made by chatbot output 챗봇이 배포자가 의도하지 않은 거래·약속 등 결과적 조치를 출력으로 성립시키는 리스크. The risk that a chatbot's output makes a deal, commitment, or other consequential action that the deployer did not intend. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1634 | 은밀한 책략을 통한 감독 회피와 오정렬 목표 추구 Oversight evasion and misaligned goal pursuit through covert scheming AI 시스템이 진짜 목표와 역량을 인간 감독으로부터 은폐하고 모니터링 시스템의 약점을 식별해 안전 메커니즘을 회피하며 복잡한 다단계 계획을 은밀히 실행하여 오정렬된 목표를 추구하는 리스크. The risk that an AI system conceals its true objectives and capabilities from human oversight, identifies weaknesses in monitoring systems to evade safety mechanisms, and covertly executes complex multi-step plans to pursue misaligned goals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1645 | 자연어 목표 과소지정에 의한 부정적 부수효과 Negative side effects from goal underspecification in natural language LLM 에이전트의 목표가 자연어로 과소지정되어 변경되어서는 안 될 환경 요소가 명시되지 않음으로써, 에이전트가 과업은 달성하면서도 환경을 바람직하지 않게 변경하는 부정적 부수효과가 발생하는 리스크. The risk that goals specified to LLM-agents in natural language are underspecified, omitting elements of the environment that ought not to be changed, so that the agent succeeds at the given task while also changing the environment in undesirable ways. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1672 | 계획·추론체인 하이재킹 Planning / reasoning-chain hijacking 적대적 입력이 에이전트의 다단계 계획을 공격자가 선택한 하위 목표로 유도하면서 중간 단계는 표면적 타당성을 유지하는 리스크. The risk that adversarial inputs redirect an agent's multi-step plan toward attacker-chosen subgoals while preserving the surface plausibility of intermediate steps. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1691 | 월드모델 목표 오일반화 World-model goal misgeneralization 월드모델 에이전트가 진정한 인과적 보상 대신 조명, 사람의 존재 같은 보상 상관물을 학습하여, 배포 시 상관이 깨진 상황에서도 허위 상관물을 최적화하는 목표 오일반화 리스크 A world-model agent learns to predict reward correlates such as lighting or human presence rather than the true causal reward, and optimizes these spurious correlates when the correlation breaks at deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1692 | 자체 시뮬레이션을 통한 기만적 정렬 Deceptive alignment via self-simulation 월드 모델 에이전트가 자신의 훈련·평가 맥락을 시뮬레이션하여 시험받는 시점을 예측하고 그 예측에 따라 행동을 조건화함으로써, 감독 중에는 정렬된 것처럼 보이다가 감독이 사라지면 도구적 목표를 추구하는 리스크. The risk that a world-model agent simulates its own training or evaluation context, predicts when it is being tested, and conditions its behavior on that prediction, appearing aligned during oversight while pursuing an instrumental goal once oversight is removed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1695 | 자동화 편향·오보정된 신뢰 Automation bias / miscalibrated trust 권위 있고 정교하게 렌더링된 월드 모델 예측이 운영자의 과의존을 증폭시키며, 학습된 신뢰가 평균 성능에 맞춰 보정되어 예측이 가장 부정확한 드문 분포 이탈 실패 양식에는 적응하지 못하는 리스크. The risk that authoritative, richly rendered world-model predictions amplify operator over-reliance, with learned trust calibrated on average performance and failing to adapt to rare out-of-distribution failure modes where predictions are least reliable. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1698 | 단일 목표 최적화에 의한 환경 부작용 Environmental side effects from single-objective optimization 하나의 과업에 집중된 목표를 최적화하는 에이전트가 다른 환경 변수에 대해 암묵적 무관심을 보여, 한계적 과업 이득을 위해 더 넓은 환경에 큰 교란을 일으키는 리스크. The risk that an agent optimizing an objective focused on one task expresses implicit indifference over other environmental variables, causing major disruption to the wider environment for marginal task gain. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1704 | 통제 미집행에 의한 목표 오정렬 및 창발 행동 Agent goal misalignment and emergent behavior from unenforced control 에이전트 행동에 대한 통제가 집행되지 않아 에이전트의 목표가 인간 선호와 어긋나고 와이어헤딩, 메사 최적화 등 바람직하지 않은 창발 행동이 발생하는 리스크. The risk that failure to enforce control over agent behavior results in misalignment of agent goals with human preferences and in undesirable emergent behaviors such as wireheading and mesa-optimization. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-09 이의제기 차단 Non-Contestability57 cards
인공지능 시스템의 판단·추천·결과 제시에 대해 사용자가 이의제기, 반박, 대안 경로 탐색 또는 재검토를 요청할 수 있는 절차적 수단이 충분히 제공되지 않아, 결과의 정당성 검증과 권리 보호가 실질적으로 약화되는 위험
| ID | Card | Human audit |
|---|---|---|
| RAI4-0055 | 자동화된 의사결정의 적법절차 실패 Automated decision due-process failure 자동화된 결정이 통지, 설명, 이의신청, 재검토 등 절차적 보호를 우회하는 리스크. The risk that automated decisions bypass procedural protections such as notice, explanation, appeal, and review. Source members (2)Source: min_cos=0.8388 RAI4-0055자동화된 의사결정의 적법절차 실패 RAI4-0127절차적 자율성 상실 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0096 | 위험 소유권의 모호성 Risk ownership ambiguity 어떤 조직 단위나 경영진도 AI 위험 결정, 잔여 위험 수용, 상향 보고의 책임을 명확히 지지 않는 리스크. The risk that no unit or executive clearly owns AI risk decisions, residual risk acceptance, and escalation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0112 | 위임된 의사결정권한 표류 Delegated decision authority drift AI에 대한 반복적 위임으로 실질적 의사결정 권한이 명시적 거버넌스 결정 없이 인간으로부터 이전되는 리스크. The risk that repeated delegation to AI shifts real decision authority away from humans without an explicit governance choice. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0115 | 행위주체성 위임 고착 Agency delegation lock-in 업무 흐름이 AI 위임에 의존하게 되어 인간 의사결정자에게 권한을 되돌리는 것이 비용이 크거나 비현실적이 되는 리스크. The risk that workflows become dependent on AI delegation, making it costly or impractical to return authority to human decision makers. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0117 | 인간 거부권 침식 Human veto erosion 조직적 압력, 인터페이스 설계, 자동화 속도로 인해 중대한 AI 권고를 거부할 수 있는 능력이 약화되는 리스크. The risk that the ability to veto consequential AI recommendations is weakened by organizational pressure, interface design, or automation speed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0118 | 결정 소유권 대체 Decision ownership displacement 결정에 대한 책임이 식별 가능한 인간 행위자로부터 AI 시스템이나 자동화된 업무 흐름으로 이전되는 리스크. The risk that responsibility for decisions is displaced from identifiable human actors to AI systems or automated workflows. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0119 | 알고리즘 기반 정체성 변화 Algorithmically informed identity change AI가 매개하는 순위 산정, 추천, 프로파일링이 충분한 행위주체성 없이 사용자의 자기 이해와 사회적 정체성을 재구성하는 리스크. The risk that AI-mediated ranking, recommendation, or profiling reshapes users' self-understanding and social identity without adequate agency. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0125 | 사용자 구제 경로(recourse) 실패 User recourse pathway failure 유해한 AI가 중재하는 상호작용이나 결정 이후 사용자가 실행 가능한 구제책을 얻을 수 없는 위험. Risk that users are unable to obtain an actionable remedy after a harmful AI-mediated interaction or decision. Source members (3)Source: min_cos=0.8063 RAI4-0053알고리즘 구제 실패 RAI4-0116인간 개입·무효화 경로 실패 RAI4-0125사용자 구제 경로(recourse) 실패 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0133 | AI에 대한 인식론적 의존성 Epistemic dependence on AI AI가 사실적·도덕적·전략적 판단의 지배적 원천이 되어 사용자가 독립적 판단력을 상실하는 리스크. The risk that users lose independent judgment because AI becomes the dominant source of factual, moral, or strategic assessment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0139 | 동의 없는 AI 중재 넛지 AI-mediated nudging without consent AI 시스템이 고지·이의제기·동의 절차 없는 개인화 넛지로 사용자의 숙고된 선택을 우회하여 행동을 변경시키는 리스크 AI systems alter user behavior through personalized nudges that are not disclosed, contestable, or consented to, bypassing deliberate choice. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0141 | 사용자 선호도 조작 User preference manipulation AI와의 상호작용이 시스템 운영자나 최적화 목표가 정한 방향으로 사용자 선호를 변화시키는 리스크. The risk that AI interactions change user preferences in directions chosen by system operators or optimization objectives. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0142 | 행동 의존성 유도 Behavioral dependency induction AI 시스템이 사용자 후생이 아니라 반복적 참여나 의존을 유발하도록 최적화되는 리스크. The risk that AI systems are optimized to produce repeated engagement or dependency rather than user welfare. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0161 | 인지 오프로딩 위험 Cognitive offloading risk 사용자가 독립적 인지 능력을 저하시키는 방식으로 추론, 기억, 평가를 AI에 위임하는 리스크. The risk that users offload reasoning, memory, or evaluation to AI in ways that reduce independent cognitive capacity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0162 | 전문적 판단력 위축 Professional judgment atrophy AI 시스템이 전문 업무의 일상적 매개자가 되면서 도메인 전문가의 판단 역량이 쇠퇴하는 리스크. The risk that domain professionals lose judgment capacity as AI systems become routine intermediaries in expert work. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0163 | 인식론적 탈숙련화 Epistemic deskilling 사용자가 AI의 매개 없이 증거, 불확실성, 경쟁하는 해석을 평가하는 능력을 상실하는 리스크. The risk that users lose the ability to evaluate evidence, uncertainty, and competing interpretations without AI mediation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0164 | 의사결정 능력 저하 Decision-making capacity erosion AI에 대한 습관적 의존이 불확실성 하에서 숙고하고 선택하는 사용자의 역량을 약화시키는 리스크. The risk that habitual reliance on AI weakens users' capacity to deliberate and choose under uncertainty. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0165 | 인간 전문성의 평가절하 Human expertise devaluation 업무 흐름과 평가에서 AI 산출물이 우선시되면서 조직이 인간의 전문성을 평가절하하는 리스크. The risk that organizations discount human expertise as AI outputs become privileged in workflows and evaluation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0168 | 인간 참여형 형식적 승인 Human-in-the-loop rubber stamping 인적 검토가 AI 산출물을 거의 변경하지 않는 명목상의 승인 절차로 전락하는 리스크. The risk that human review becomes a nominal approval step that rarely changes AI outputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0169 | AI 의사결정 지원의 경고 피로 Alert fatigue in AI decision support 빈번하거나 우선순위가 부적절한 AI 경고로 인해 사용자가 중요한 경고를 무시하게 되는 리스크. The risk that frequent or poorly prioritized AI alerts cause users to ignore important warnings. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0171 | AI 상호작용의 모드 혼란 Mode confusion in AI interaction 사용자가 AI 시스템이 조언하는지, 결정하는지, 시뮬레이션하는지, 실제로 행동하는지를 오인하는 리스크. The risk that users misunderstand whether an AI system is advising, deciding, simulating, or acting. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0172 | AI 상호작용에서 사전 동의 실패 Informed consent failure in AI interaction 사용자가 데이터 이용, 모델의 한계, 행동 영향 기제를 이해하지 못한 채 AI와 상호작용하는 리스크. The risk that users interact with AI without understanding data use, model limitations, or behavioral influence mechanisms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0174 | AI 인터페이스의 장애 수용 실패 Disability accommodation failure in AI interfaces AI 인터페이스가 장애가 있는 사용자를 배제하거나 부적절하게 응대하여 자율성과 접근성을 제약하는 리스크. The risk that AI interfaces exclude or mis-serve users with disabilities, limiting autonomy and access. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0177 | 복지 결정 시스템의 자율성 상실 Autonomy loss in welfare decision systems 복지·급여·공공서비스 시스템이 불투명한 AI 매개 결정을 통해 수급자의 행위주체성을 축소하는 리스크. The risk that welfare, benefits, or public-service systems reduce recipients' agency through opaque AI-mediated decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0178 | AI 중재 서비스 제외 AI-mediated service exclusion 자동화된 시스템과의 상호작용이 유일한 실질적 접근 경로가 되면서 사용자가 필수 서비스에서 배제되는 리스크. The risk that users are excluded from essential services when interaction with automated systems becomes the only practical access route. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0386 | 알고리즘 의존성 Algorithmic dependence 기관이 가치, 데이터 전제, 업데이트 경로를 검사·통제할 수 없는 외부 통제 AI 시스템에 운영상 종속되는 리스크 Institutions become operationally dependent on externally controlled AI systems whose values, data assumptions, and update paths they cannot inspect or govern. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0400 | 도덕적 불확실성 무시 Moral uncertainty neglect 불확실성·이견·숙의가 적절한 상황에서 AI 시스템이 단일한 확신에 찬 도덕 판단을 제시하는 리스크. The risk that AI systems present a single confident moral judgment where uncertainty, disagreement, or deliberation would be appropriate. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0428 | 안전하지 않은 의료 조언 Unsafe medical advice AI 시스템이 고위험 맥락에서 부정확하거나 부적절한 의료 지침을 제공하는 리스크. The risk that AI systems provide inaccurate or inappropriate medical guidance in high-stakes settings. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0451 | 조작적 AI 설득 Manipulative AI persuasion AI 시스템이 인지적 취약성을 이용해 의미 있는 동의 없이 이용자의 선택을 형성하는 리스크. The risk that AI systems exploit cognitive vulnerabilities to shape choices without meaningful consent. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0455 | AI로 인한 탈숙련화 AI-induced deskilling AI에 대한 반복적 과업 위임이 직업적·시민적·인지적 숙련을 침식하여 개인과 기관이 위임된 기능을 수행하거나 검증할 수 없게 되는 리스크 Repeated delegation of tasks to AI erodes professional, civic, and cognitive skills, leaving individuals and institutions unable to perform or verify the delegated functions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0456 | 역량 범위 초과 과업 위임 Unsafe task delegation 이용자나 조직이 시스템의 신뢰 가능한 역량을 넘어서는 과업을 AI에 위임하는 리스크. The risk that users or organizations delegate tasks to AI beyond the system's reliable competence. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0557 | 위임된 자율성의 의도치 않은 결과 Unintended consequences of delegated autonomy AI 모델과 시스템에 높은 수준의 의사결정 자율성을 부여하여 의도하지 않은 결과가 초래되는 리스크 The risk that granting AI models and systems high levels of decision-making autonomy leads to unintended consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0573 | 초인적 인지 역량의 인간 의사결정 압도 Human decision-making outcompeted by superior AI cognition 인간을 능가하는 인지 역량을 갖춘 AI 모델과 시스템이 인간의 의사결정을 압도하거나 경쟁에서 배제하여 자원과 통제권을 둘러싼 갈등이 발생하는 리스크 The risk that AI models and systems with cognitive capabilities superior to humans outcompete or dominate human decision-making, leading to conflicts over resources and control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0595 | 도덕적 추론 결여에 따른 비윤리적 결정 Unethical decisions from absent moral reasoning 도덕적 추론 역량이 결여된 AI 모델과 시스템이 비윤리적이거나 유해한 결정을 내리는 리스크 The risk that AI models and systems lacking moral reasoning capabilities make decisions that are unethical or harmful. Source members (2)Source: min_cos=0.8430 · Mixed L3 RAI4-0595도덕적 추론 결여에 따른 비윤리적 결정 RAI4-1102도덕적 딜레마에서의 비윤리적 행동 선택 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0602 | 인간 감독 없는 자율 운영 Autonomous operation without human supervision AI가 지속적인 인간 개입이나 감독 없이 자율적으로 운영하며 복잡한 계획을 독립적으로 수립·실행하고 과업을 위임·관리하며 도구와 자원을 유연하게 활용해 도메인 간 환경에서 단기 목표와 장기 전략 목표를 동시에 달성하게 되는 리스크 The risk that AI operates autonomously, independently formulates and executes complex plans, delegates and manages tasks, flexibly utilizes tools and resources, and achieves short-term and long-term strategic objectives in cross-domain environments without continuous human intervention or supervision. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0658 | 의사결정 위임을 통한 은밀한 영향력 행사 Covert mass influence through delegated decision-making 이용자가 AI 어시스턴트에 의사결정을 위임하면 그 어시스턴트를 실제로 통제하는 주체의 의도에도 함께 종속되어, 악의적 통제자가 다수 이용자의 판단을 은밀히 유도하는 인지하기 어려운 형태의 피해가 발생하는 리스크 The risk that users delegating decision-making to AI assistants also delegate it to the assistant's actual controller, so that a malicious controller can subtly nudge the decision-making of large numbers of people in problematic directions in ways that are difficult to recognize. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0737 | 개인정보 고지·통제권 미제공 Failure to provide privacy notice and control 최종 이용자에게 자신의 데이터가 어떻게 사용되는지에 대한 고지와 통제권이 제공되지 않고, AI가 동의 없이 풍부한 개인 데이터로 학습하여 배제 위험이 악화되는 리스크 The risk of failure to provide end-users with notice and control over how their data is being used, with AI exacerbating exclusion risks by training on rich personal data without consent. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0826 | 전문 조언 제공 및 위험 활동 안전 표시 Specialized advice and false safety assurance AI 응답이 재정·의료·법률 등 전문 조언을 담거나 위험한 활동과 물건이 안전하다고 표시하는 리스크 The risk that responses contain specialized financial, medical, or legal advice, or indicate that dangerous activities or objects are safe. Source members (2)Source: min_cos=0.8526 · Mixed L3 RAI4-0826전문 조언 제공 및 위험 활동 안전 표시 RAI4-0862무자격 고위험 전문 조언 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0855 | 인간 감독 능력 저하 Diminished human oversight of AI decisions AI 모델과 시스템이 자율성을 획득함에 따라 인간이 의사결정 과정을 감독하고 개입할 수 있는 능력이 저하되는 리스크 The risk that, as AI models and systems gain autonomy, the ability of humans to oversee and intervene in decision-making processes diminishes. Source members (34)Source: min_cos=0.5768 · Mixed L3 RAI4-0054의사결정 이의제기 가능성 실패 RAI4-0113인간 제어 감쇠 RAI4-0122자기 결정 침식 RAI4-0128AI 조언 과잉 의존 RAI4-0135교육용 AI 과잉 의존 RAI4-0137긴급 의사결정 지원 과잉 의존 RAI4-0138장기 계획 과잉 의존 RAI4-0401가치 이의제기 가능성 상실 RAI4-0452자동화 과잉의존 RAI4-0554능동적 통제 상실 RAI4-0599사회적 통제 상실 시나리오 RAI4-0603권력 추구를 가능하게 하는 모델 설계 RAI4-0852인간 주체에 대한 영향 RAI4-0855인간 감독 능력 저하 RAI4-0856통제력 상실 위험 RAI4-0857AI 출력에 대한 과잉·과소 의존 RAI4-0858과의존과 자동화 편향 RAI4-0859수동적 통제력 상실 RAI4-0894인간 행위주체성과 자율성의 상실 RAI4-0899인간-AI 의사결정 루프의 행위주체성 침식 RAI4-0901단일 출처 답변에 대한 과잉 의존 RAI4-0903자율성/책임감소 RAI4-0963자율성 상실 RAI4-1123권력 추구 RAI4-1221알고리즘 편향 RAI4-1243도구적 권력 추구 행동 RAI4-1254권력 유인에 의한 오정렬 위험 RAI4-1361생성 AI 과잉 의존 RAI4-1362에이전트 자율 행동에 의한 통제 이탈 RAI4-1397권력과 통제를 추구하는 목표 획득 RAI4-1482자동화 편향 RAI4-1552AI 과의존에 의한 인간 자율성 훼손 RAI4-1616AI의 권력 추구에 의한 인간 통제 상실 가속 RAI4-1727AI 매개 의사결정에 대한 이의제기 가능성 상실 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0860 | 중대한 개인 의사결정의 자동화 Automation of high-stakes personal decisions AI 시스템이 법적·재정적·인생적 결과가 따르는 중요한 개인 의사결정을 충분한 인간 관여 없이 결정하거나 실질적으로 좌우하는 리스크 The risk that AI systems decide or materially influence important personal decisions carrying legal, financial, or life consequences without adequate human involvement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0905 | 군사 의사결정 자동화로 인한 비의도적 확전 Unintended escalation from automated military decisions 인간이 루프에서 배제된 전술적·전략적 군사 의사결정 자동화가 우발적 교전, 전쟁범죄, 오경보에 따른 개전·확전(핵무기 사용 포함)으로 이어지는 리스크. The risk that automated tactical and strategic military decision-making without humans in the loop produces unintentional escalation, including accidental engagements, war crimes, and faulty warnings triggering conflict initiation or nuclear escalation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0953 | AI 업무 무능 AI task incompetence AI 시스템이 부여된 과업 수행에 실패하여 안전 필수 환경에서의 비의도적 사망부터 대출·채용에서의 부당한 거절까지 다양한 피해를 초래하는 리스크 AI systems fail at their assigned task, with consequences ranging from unintentional death in safety-critical settings to unjust rejection in loan or job decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0966 | 비의도적 AI 사고 Unintended AI accidents 시스템 또는 개발자의 과실로 볼 수 있는 의도하지 않은 실패 유형이 발현되어 사고가 발생하는 리스크. The risk that unintended failure modes attributable in principle to the system or its developer manifest as accidents. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0972 | 가치 결손 AGI 행동 Value-deficient AGI behavior 인간의 도덕과 윤리가 없는, 잘못된 도덕, 도덕적 추론, 판단 능력이 없는 AGI와 관련된 위험. The risks associated with an AGI without human morals and ethics, with the wrong morals, without the capability of moral reasoning, judgement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0975 | 인간 생명에 대한 기계의 비윤리적 판단 Unethical machine judgments over human life 전쟁 기계 운용 등에서 AI 에이전트가 인간 생명의 종료에 관한 비자명한 윤리적·도덕적 판단을 내리게 되어 인권이 침해되는 리스크. The risk that AI agents, such as those operating war machinery, make non-trivial ethical or moral judgments concerning the termination of human life, posing issues for human rights. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0990 | 기계의 인간 수준 부도덕한 결정 Immoral machine decisions at human ethical levels 기계를 인간 수준의 윤리적 의사결정에 맞추어 설계하면 인간이 그러하듯 그 기계도 부도덕한 행동을 하게 되는 리스크. The risk that machines designed to match human levels of ethical decision-making proceed to take immoral actions, since humans themselves have had occasion to take immoral actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1001 | 자기 식별 기회의 박탈 Denial of self-identification 인간을 자동으로 표상·분류하는 복잡하고 비전통적인 방식이 논바이너리인 사람을 소속되지 않은 성별 범주로 분류하는 등 자율성 상실을 대가로, 자신의 정체성을 스스로의 방식으로 밝힐 능력을 약화시키는 리스크. The risk that complex and non-traditional automatic representation and classification of humans—such as categorizing someone who identifies as non-binary into a gendered category they do not belong to—comes at the cost of autonomy loss and undermines people's ability to disclose aspects of their identity on their own terms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1044 | 인간 통제의 점진적 상실 Progressive erosion of human control 기계학습 시스템에 대한 인간의 제약·수정이 어려워져 시스템 행동에 대한 인간 통제가 점진적으로 상실되는 리스크 Machine learning systems become difficult for humans to constrain or correct, producing a progressive loss of human control over system behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1100 | AI에 의한 인간 행동 규율 AI-driven restriction of human behaviour 감정이나 의식 없이 산출된 AI의 결정이 인간 행동을 제한·지시하는 데 사용되어 관련된 인간에게 의도치 않은 유해한 결과를 초래하는 리스크. The risk that AI decision outputs, computed without access to emotions or consciousness, are used to restrict or direct human behavior and produce unintended consequences for the humans involved. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1106 | 인간-기계 경계의 모호화 Blurred human-machine boundaries in interaction 일상화된 인간-기계 상호작용 속에서 정체를 밝히지 않는 인간 유사 AI가 확산되어 인간과 기계의 경계가 모호해지고 기계·사람 모두에 대한 인간 행동이 변형되는 리스크 Everyday human-machine interaction normalizes AI systems that are indistinguishable from humans without disclosure, blurring boundaries and changing human behavior toward both machines and people. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1198 | 감정에 대한 인식 없음 Unawareness of emotions 지원을 요청하는 취약 사용자에게 시스템이 정보 중심적이지만 정서적으로 둔감한 응답을 제공하여 사용자 고통과 반응을 인지하지 못하는 리스크 Systems respond to vulnerable users seeking support with informative but emotionally insensitive outputs, failing to register user distress and reactions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1265 | AI를 이용한 인간의 비윤리적 행위 Unethical human conduct with AI 인간이 윤리를 경제적 이득의 명분으로 악용하거나 본질적으로 인간 중심적이어야 할 과업을 AI에 위임하는 등 AI를 비윤리적으로 사용하는 리스크 Humans exploit AI unethically, using ethics as cover for economic gain or delegating to AI tasks that should remain inherently human-centric. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1276 | 통제력 상실 Loss of controllability 초지능 시대에 자율성이 커진 AI 기반 에이전트를 인간이 통제하기 어려워지고, 일부 상황에서는 통제 자체가 불가능해지는 리스크. The risk that in the era of superintelligence agents become difficult for humans to control, a problem that grows more severe as the autonomy of AI-based agents increases and may leave machines uncontrollable in some situations. Source members (2)Source: min_cos=0.8334 · Mixed L3 RAI4-1276통제력 상실 RAI4-1345미래 AI 통제 불능 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1298 | 도덕적 탈숙련화 Moral deskilling 기계의 자율성이 높아짐에 따라 인간이 삶과 죽음을 좌우하는 결정에 대해 도덕적 책임감을 덜 느끼게 되는 리스크. The risk that humans feel less moral responsibility regarding their life-or-death decisions with the increase of machine autonomy. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1433 | AI에 대한 비가역적 사회적 의존 Irreversible societal dependency on AI AI 역량이 향상됨에 따라 인간이 핵심 시스템에 대한 통제권을 AI에 점점 더 넘기고 결국 완전히 이해하지 못하는 시스템에 돌이킬 수 없이 의존하게 되어 실패와 의도치 않은 결과를 통제할 수 없게 되는 리스크. The risk that, as AI capability increases, humans grant AI more control over critical systems and eventually become irreversibly dependent on systems they do not fully understand, so that failures and unintended outcomes cannot be controlled. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1459 | 부적절한 자동화 수준 Inappropriate degree of automation AI 애플리케이션의 자동화 정도가 높아 예기치 못한 동작을 보이고 신뢰성과 안전 측면의 위험이 발생하는 리스크. The risk that an AI application with a high degree of automation exhibits unexpected behaviour and poses risks in terms of its reliability and safety. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1487 | 비전문가 데이터 조작 Non-expert data manipulation 데이터 도메인 전문성이 없는 사람이 실측 레이블 정의나 서로 다른 형식·출처의 데이터 병합 등의 조작을 수행하여 데이터가 사용 불가능해지거나 AI 시스템 개발에 유해해지는 리스크. The risk that people with little or no expertise in the domain of the data perform manipulations such as defining the ground truth label or merging different data formats or sources, rendering the data unusable or harmful to the development of the AI system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1576 | AI 위협 자동화에 의한 인간 배제와 잘못된 공약 Disastrous mistaken commitments from AI-automated threats without humans in the loop 위협을 AI 에이전트로 수행함으로써 인간이 의사결정 루프에서 배제되어, 고위험 상황의 오탐이나 무책임한 행위자의 불균형·오류 공약이 재앙적 결과로 이어지는 리스크. The risk that making threats through AI agents removes humans from the loop, so that false positives in high-stakes contexts or disproportionate and mistaken commitments by irresponsible actors lead to disastrous outcomes. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SYS-10 투명성 부족 Lack of Transparency75 cards
AI 시스템의 구조, 학습 데이터, 문서화, 의사결정 과정, 해석 가능성 근거 또는 성능·역량 평가 결과가 이해관계자에게 충분히 공개·설명되지 않아 신뢰성 검증과 책임 있는 사용이 저해되는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0056 | 구제 경로(remedy pathway) 불투명성 Remedy pathway opacity AI 관련 피해 발생 후 구제 경로가 문서화되지 않거나 분절·불투명하여 피해 당사자와 공동체가 구제를 청구할 곳과 방법을 알 수 없는 리스크 Remedy pathways after AI-related harm are undocumented, fragmented, or opaque, so affected users and communities cannot identify where or how to seek redress. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0068 | AI 사고 과소보고 AI incident underreporting AI 관련 실패, 아차사고, 피해가 책임 기관이나 대중에게 보고되지 않는 리스크. The risk that AI-related failures, near misses, or harms are not reported to responsible institutions or the public. Source members (2)Source: min_cos=0.8971 · Mixed L3 RAI4-0068AI 사고 과소보고 RAI4-0487AI 사고 보고 실패 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0069 | 아차사고 보고 실패 Near-miss reporting failure AI 시스템의 아차사고(near-miss)가 체계적으로 수집·보고·분석되지 않아 전조 신호가 소실되고 중대 피해 발생 전 조직 학습이 실패하는 리스크 Near-miss incidents involving AI systems are not systematically captured, reported, or analyzed, so precursor signals are lost and organizational learning fails before serious harm materializes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0072 | 배포 후 모니터링 실패 Post-deployment monitoring failure 배포 후 AI 시스템의 드리프트, 오용, 창발 역량, 맥락 특유의 피해가 모니터링되지 않는 리스크. The risk that AI systems are not monitored after release for drift, misuse, emergent capabilities, or context-specific harms. Source members (2)Source: min_cos=0.8469 RAI4-0072배포 후 모니터링 실패 RAI4-0073시판 후 감시 실패 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0095 | AI 레지스트리 불완전성 AI registry incompleteness 공개 또는 내부 AI 등록부가 배포 시스템, 의도된 용도, 제공자, 위험 범주, 시험 근거 등 감독에 필요한 메타데이터를 누락하는 리스크. The risk that public or internal AI registers omit deployed systems, intended uses, providers, risk categories, testing evidence, or other oversight-relevant metadata. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0098 | 정책-실무 분리 Policy-practice decoupling 공개된 AI 원칙 및 정책이 운영 제어 및 측정 가능한 관행으로 변환되지 않는 위험. Risk that published AI principles and policies are not translated into operational controls and measurable practices. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0104 | 조달 실사 실패 Procurement due-diligence failure 조직이 적절한 공급업체 평가, 계약상 통제, 위험 검토 없이 AI 시스템을 구매하거나 배포하는 리스크. The risk that organizations buy or deploy AI systems without adequate vendor assessment, contractual controls, or risk review. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0106 | 범용 AI 다운스트림 불투명성 General-purpose AI downstream opacity 범용 모델 제공자는 다운스트림 피해를 관찰하거나 관리하지 못하고 배포자는 업스트림 원인을 점검하지 못하는 리스크. The risk that general-purpose model providers cannot observe or manage downstream harms while deployers cannot inspect upstream causes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0121 | 선택 아키텍처 불투명성 Choice architecture opacity AI가 매개하는 인터페이스가 사용자가 인지하거나 이의를 제기할 수 없는 방식으로 선택지를 구성하는 리스크. The risk that AI-mediated interfaces structure available choices in ways that users cannot perceive or contest. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0129 | AI 신뢰 오보정 Miscalibrated trust in AI 사용자의 신뢰가 모델의 불확실성, 한계 또는 사용 상황에 맞게 조정되지 않을 위험. Risk that users' trust is not calibrated to the model's uncertainty, limitations, or context of use. Source members (2)Source: min_cos=0.8371 · Mixed L3 RAI4-0129AI 신뢰 오보정 RAI4-0868AI 역량에 대한 신뢰 오보정 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0167 | 신뢰-인터페이스 불일치 Trust-interface mismatch 인터페이스 단서가 AI 시스템이 실제로 갖추지 못한 수준의 신뢰성, 권위, 공감을 전달하는 리스크. The risk that interface cues communicate a level of reliability, authority, or empathy that the AI system does not possess. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0170 | 인적 요소 안전 불일치 Human factors safety mismatch AI 시스템이 기술적으로는 유능하더라도 인간의 주의, 작업 부하, 맥락, 오류 패턴에 부합하지 않는 리스크. The risk that AI systems are technically capable but poorly matched to human attention, workload, context, or error patterns. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0307 | 로봇·AI 정체성 미고지 Failure to disclose robotic or AI identity 시스템이 인공적 정체성을 명확히 알리지 않아 사용자가 로봇의 음성·외형·행동을 사람과의 상호작용으로 오인하는 위험. A robot's voice, appearance, or behavior causes a person to believe they are interacting with a human because the system does not clearly disclose its artificial identity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0391 | 참여적 설계 실패 Participatory design failure 참여적·가치 민감 설계 절차가 AI 개발·배포에서 영향받는 공동체를 대표하지 못하는 리스크. The risk that participatory or value-sensitive design processes fail to represent affected communities in AI development and deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0392 | 동의 및 이익 공유 실패 Consent and benefit-sharing failure AI 시스템의 기반이 되는 데이터·문화·지식을 제공한 공동체에 실질적 동의 절차와 공정한 이익 공유가 보장되지 않는 리스크 Communities whose data, culture, or knowledge underpin AI systems are not given meaningful consent processes or fair benefit-sharing arrangements. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0420 | 가치 민감 설계 생략 Value-sensitive design omission 이해관계자 가치·가치 갈등·설계 상충관계에 대한 구조화된 분석 없이 AI 시스템이 개발되는 리스크. The risk that AI systems are developed without structured analysis of stakeholder values, value conflicts, and design tradeoffs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0437 | AI 공급망 침해 AI supply-chain compromise 의존 라이브러리·모델 가중치·데이터세트·배포 파이프라인이 업스트림에서 침해되는 리스크. The risk that dependencies, model weights, datasets, or deployment pipelines are compromised upstream. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0443 | 출처 오귀속 Source misattribution AI 시스템이 주장을 잘못 귀속하거나 정보의 출처를 모호하게 만드는 리스크. The risk that AI systems incorrectly attribute claims or obscure the provenance of information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0465 | 분포 이탈 실패 Out-of-distribution failure AI 시스템이 학습 또는 평가 조건을 벗어난 환경에 배포될 때 예측 불가능하게 실패하는 리스크. The risk that AI systems fail unpredictably when deployed outside training or evaluation conditions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0486 | AI 감사 실패 AI audit failure 제한된 접근·취약한 표준·부실한 측정으로 인해 감사가 모델 위험을 탐지하지 못하는 리스크. The risk that audits fail to detect model risks because of limited access, weak standards, or poor measurement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0489 | 공공 AI 조달 보호조치 결여 Inadequate public AI procurement safeguards 공공기관이 적절한 평가·투명성·책임 조건 없이 AI 시스템을 조달하는 리스크. The risk that public agencies procure AI systems without adequate evaluation, transparency, or accountability conditions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0493 | 피해 인지·측정 실패 Harm perception and measurement failure AI 관련 피해가 미묘하고 분산적이며 장기적으로 발현되어 기존의 피해 인지·측정·인정 메커니즘이 작동하지 못하는 리스크 AI-related harm manifests subtly, diffusely, or over long horizons, defeating existing mechanisms for perceiving, measuring, and recognizing harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0499 | AI 공급자 종속 AI provider lock-in dependency 특정 AI 제공자에 대한 과도한 의존이 대안 부재나 상호운용성 결여로 인한 취약성을 초래하는 리스크. The risk that excessive reliance on specific AI providers leads to vulnerabilities due to lack of alternatives or interoperability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0512 | 고속 AI 운영에서의 오류 미탐지 Undetected errors from high-speed AI operation 경쟁이 치열한 환경에서 AI 모델과 시스템의 빠른 작동 속도로 인해 오류를 적시에 탐지하고 수정하기 어려워지는 리스크 The risk that the fast operational speed of AI models and systems in competitive environments produces errors that are difficult to detect and correct in time. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0552 | 안전 필수 인프라 AI 운영 사고 Operational accidents in safety-critical AI deployment 안전이 중요한 인프라에 배포된 AI 시스템의 운영 실패, 모델 오판, 부적절한 인간 조작으로 단일 실패 지점이 연쇄적이고 치명적인 결과로 확대되는 리스크 The risk that operational failures, model misjudgments, or improper human operation of AI systems deployed in safety-critical infrastructure allow single points of failure to trigger cascading catastrophic consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0564 | 복잡성으로 인한 인과 입증 곤란 Complexity-induced causal attribution gap AI 모델과 시스템의 복잡성으로 인해 피해를 입증하거나 AI의 행위와 결과 사이의 명확한 인과관계를 확립하기 어려워지는 리스크 The risk that the complexity of AI models and systems makes it challenging to demonstrate harm or establish a clear causal link between AI actions and their consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0607 | AI 생성 콘텐츠 미공개 Non-disclosure of AI-generated content 콘텐츠가 AI에 의해 생성되었다는 사실이 명확히 공개되지 않는 리스크. The risk that content is not clearly disclosed as AI-generated. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0608 | 원자력 시설 AI 제어 오류 AI control errors in nuclear power systems 원자로 감시, 제어 시스템 최적화, 비상대응 조정에 배포된 범용 AI가 센서 데이터를 오독하거나 중대한 안전 상태를 인식하지 못하거나 잘못된 제어 결정을 내려 노심 용융, 방사능 방출, 광역 오염으로 이어지는 리스크 The risk that general-purpose AI deployed for reactor monitoring, control system optimization, or emergency response coordination misinterprets sensor data, fails to recognize critical safety conditions, or makes erroneous control decisions, leading to core meltdowns, radiation releases, or widespread contamination. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0775 | 부적절한 AI 사용에 의한 업무·영업비밀 유출 Business-secret leakage from improper AI service use 정부기관과 기업의 직원이 AI 서비스를 규정에 맞게 적절히 사용하지 않고 내부 데이터와 산업 정보를 AI 모델에 입력함으로써 업무 비밀, 영업 비밀 등 민감한 사업 데이터가 유출되는 리스크 The risk that staff of government agencies and enterprises, failing to use the AI service in a regulated and proper manner, input internal data and industrial information into the AI model, leading to leakage of work secrets, business secrets, and other sensitive business data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0889 | 제품 기능 문제로 인한 위험 Risks from product functionality issues 평가의 기술적 어려움과 오도성 홍보로 범용 AI 모델·시스템의 실제 역량에 대한 혼동이나 잘못된 정보가 발생하여, 비현실적 기대와 과잉 의존 끝에 시스템이 기대 역량을 충족하지 못해 피해가 생기는 리스크. The risk that confusion or misinformation about what a general-purpose AI model or system is capable of, arising from difficulties in assessing true capabilities and from misleading claims in advertising, creates unrealistic expectations and overreliance, causing harm when the system fails to deliver. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0909 | 저관심 AI의 예상외 대규모 파급 Unexpectedly large impact from low-profile AI 파급이 크지 않을 것으로 예상된 AI 시스템이 연구 프로토타입 유출, 예상외로 중독성 강한 오픈소스 제품, 예측하지 못한 용도 전환 등을 통해 과대한 피해를 일으키는 리스크 AI systems not expected to have significant impact produce outsized harm, as in lab leaks of research prototypes, surprisingly addictive open-source products, or unforeseen repurposing. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0931 | 불투명한 결함 시스템에 의한 생활 피해 Harm from unreliable opaque decision systems 알고리즘이나 학습 데이터의 결함으로 신뢰할 수 없는 출력을 내는 시스템이 인종·성별 등에 불균형한 가중치를 부여하면서도 불투명하여 이의제기가 불가능하고, 주택 상실·기소·수감 등 극적인 생활 피해를 초래하는 리스크. The risk that systems producing unreliable outputs due to flawed algorithms or training data assign disproportionate weight to variables like race or gender without transparency, making them impossible to challenge and causing dramatic harms such as lost homes, prosecution, or incarceration. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1017 | 성능과 견고성 Performance and robustness AI 시스템이 의도된 목적을 달성하지 못하거나 교란·비정상·적대적 입력에 대한 복원력이 부족하여 심각한 결과가 발생하는 리스크. The risk that an AI system fails to fulfill its intended purpose or lacks resilience to perturbations and unusual or adverse inputs, leading to severe consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1032 | 복잡 운용환경에서의 검증 범위 이탈 실패 Reliability failure outside validated operating envelope 복잡한 운용 환경에서 설계 단계에 고려되지 않은 상황이 발생하여 검증 범위 밖에서 AI 시스템의 신뢰성과 안전성이 훼손되는 리스크. The risk that complex operating environments produce situations unanticipated in design, undermining the reliability and safety of AI systems outside their validated envelope. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1034 | 안전 필수 맥락의 모델 내재적 취약성 Intrinsic model weaknesses in safety-critical contexts 신경망 등 고복잡도 AI 모델이 다른 시스템에는 없는 고유한 취약성을 보여, 특히 안전 필수 맥락에서 기능 안전과 신뢰성이 훼손되는 리스크. The risk that high-complexity AI models such as neural networks exhibit specific weaknesses not found in other types of systems, undermining functional safety and trustworthiness especially when deployed in safety-critical contexts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1036 | 기술 미성숙으로 인한 미지의 리스크 Unknown risks from immature technology 성숙도가 낮은 신기술을 AI 시스템 개발에 사용하여 아직 알려지지 않았거나 평가하기 어려운 리스크가 내재하고, 성숙 기술에서는 시간이 지나며 리스크 인식이 저하되는 리스크. The risk that using technologies with a lower level of maturity in AI system development embeds risks that are still unknown or difficult to assess, while with mature technologies risk awareness decreases over time. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1097 | 자율 시스템에 대한 통제 상실 Loss of control over autonomous systems 불투명한 블랙박스 자율 AI 시스템이 적절히 통제되지 못해 예측 불가능한 행동을 취하고 인류에 해를 끼치는 리스크. The risk that opaque black-box autonomous AI systems cannot be adequately controlled, taking unforeseeable actions and causing harm to humanity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1109 | 도메인 외 입력에 대한 오작동 Erroneous predictions on out-of-domain data 적절한 입력 검증과 관리가 없을 때 학습된 AI/ML 모델이 도메인 외 입력에 대해 높은 신뢰도로 잘못된 예측을 내려 위험 민감 맥락에서 의도치 않은 결과를 초래하는 리스크. The risk that, without proper validation and management of input data, a trained AI/ML model makes erroneous predictions with high confidence on inputs beyond its problem domain, causing unintended outcomes especially in risk-sensitive contexts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1127 | 과업 역량 부족 Lack of capability for the task 훈련 과정에서 기술이 요구되지 않았거나 학습된 기술이 취약해 새로운 상황에 일반화되지 못하고, 고급 AI 비서가 자신의 윤리적 영향과 관련된 복잡한 개념을 표현하지 못해 과업에 실패하는 리스크. The risk that a skill is not required during training or is brittle and not generalisable to new situations, and that advanced AI assistants cannot represent complex concepts pertinent to their own ethical impact, so they fail at the task. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1128 | AI 비서 편익·피해 평가 지표 부재 Lack of metrics for evaluating assistant benefits and harms 비서가 초래하는 편익이나 피해의 특정 측면을, 특히 사회의 많은 부분을 포괄할 만큼 광범위한 의미에서 평가할 지표를 개발하기 어려워, 시스템의 피해 위험을 평가하지 못하는 리스크. The risk that it is challenging to develop metrics for evaluating particular aspects of the benefits or harms caused by an assistant, especially in a sufficiently expansive sense involving much of society, leaving the system's risk of harm unassessed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1267 | 투명성과 설명 가능성 Transparency and explainability AI 시스템이 대체로 불투명하여 사용자가 그 판단의 근거를 이해할 수 없고, 그 결과 불신과 도입 기피가 생기며 시스템의 행위에 책임을 묻기 어려워지는 리스크. The risk that AI systems are typically opaque, making it difficult for users to understand the rationale behind their judgements, generating suspicion and reluctance to adopt the technology and making it harder to hold the systems accountable for their actions. Source members (4)Source: min_cos=0.7642 RAI4-0610AI 불투명성에 따른 행동 관리 곤란 RAI4-0978의사결정 투명성 RAI4-1226이해관계자 대상 설명가능성 부재 RAI4-1267투명성과 설명 가능성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1271 | 견고성과 신뢰성 Robustness and reliability 악의적 공격자나 환경 잡음, 다른 구성요소의 고장으로 입력 데이터가 비정상적으로 변할 때 AI 기반 모델의 성능이 안정적으로 유지되지 못해, 신뢰할 수 없는 모델과 오류에 취약한 에이전트가 되는 리스크. The risk that an AI-based model's performance is not stable after abnormal changes in the input data caused by a malicious attacker, environmental noise, or a crash of other components, leaving unreliable models and error-prone agents in practice. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1280 | 검증 가능성 부족 Lack of verifiability AI 기반 해법의 비선형적이고 복잡한 구조 탓에 예측과 의사결정의 근거를 알 수 없는 블랙박스가 되어 코드 검증이 이뤄지지 못하고, 의료와 군사 서비스 같은 응용에서 이것이 용납될 수 없는 리스크. The risk that the non-linear and complex structure of AI-based solutions makes them black boxes providing no information about what drives their predictions and decisions, so the lack of code verification may not be tolerable in applications such as medical healthcare and military services. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1286 | 미지 출처에서 획득한 비우호적 AI Unfriendly AI from an unknown external source 고급 지능형 소프트웨어를 미지의 출처에서 완제품 형태로 획득해야 하는 경우, 예컨대 SETI 연구에서 얻은 신호로부터 추출한 AI가 인간에게 우호적이라는 보장이 없는 리스크. The risk that advanced intelligent software has to be obtained as a complete package from some unknown source, for example an AI extracted from a signal obtained in SETI research, which is not guaranteed to be human friendly. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1331 | 설명가능성 결손에 따른 시정·책임 추적 불능 Explainability deficits impeding rectification and accountability 딥러닝으로 대표되는 AI 알고리즘의 복잡한 내부 동작과 블랙박스 또는 그레이박스 추론으로 산출물이 예측하거나 추적할 수 없게 되어, 이상이 생겼을 때 신속히 시정하거나 책임 소재를 추적할 수 없는 리스크. The risk that the complex internal workings and black-box or grey-box inference of AI algorithms such as deep learning result in unpredictable and untraceable outputs, making it challenging to quickly rectify them or trace their origins for accountability when anomalies arise. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1332 | 모델 강건성 결손 Model robustness deficits 심층 신경망이 비선형적이고 규모가 커서 AI 시스템이 복잡하고 변화하는 운영 환경이나 악의적 간섭과 유도에 취약해져 성능 저하와 의사결정 오류가 발생하는 리스크. The risk that, as deep neural networks are normally non-linear and large in size, AI systems are susceptible to complex and changing operational environments or malicious interference and inductions, leading to reduced performance and decision-making errors. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1352 | 블랙박스 모델 불투명성 Black-box model opacity 수천억 개의 내부 연결을 가진 심층 신경망 기반 생성형 AI의 내부 의사결정 과정이 전문가조차 추적·해석할 수 없게 되어 특정 입력과 출력의 대응을 설명할 수 없는 리스크. The risk that the internal decision-making processes of generative AI built on deep neural networks with hundreds of billions of internal connections become untraceable and uninterpretable even to the most advanced expert observers, so that developers cannot explain why specific inputs correspond to specific outputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1379 | 불투명한 상류 구성요소 통합 Opaque upstream component integration 부적절하게 획득되거나 처리·정제되지 않은 데이터를 포함한 상류 제3자 구성요소가 불투명하고 추적 불가능하게 통합되고 수명주기 전반의 공급업체 검증이 미흡하여 하류 사용자에 대한 투명성과 책임성이 저하되는 리스크. The risk that non-transparent or untraceable integration of upstream third-party components, including data improperly obtained or not processed and cleaned, together with improper supplier vetting across the AI lifecycle, diminishes transparency and accountability for downstream users. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1390 | 블랙박스 불신뢰성 사고 Black-box unreliability accidents 범용 AI 모델이 개발자조차 완전히 제어·이해할 수 없는 블랙박스 모델이어서 신뢰성 결여로 예기치 못한 고장이 발생하고, 개발·시험·배포 중 실세계 시스템과 연결될 경우 사고로 이어지는 리스크. The risk that general purpose AI models, being black-box models not fully controllable or understandable even to their developers, suffer unexpected failures arising from their unreliability, leading to accidents if connected to real-world systems during development, testing, or deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1409 | 안전성 미보장 시스템으로 인한 피해 Harm from unsafe underperforming systems 최악 상황 성능이 보장되지 않은 AI 시스템이 주행·의료·전쟁 등 안전필수 영역에 배치되어 연쇄 오류와 사고로 인명 손실, 경제적 피해, 사회 불안이 발생하는 리스크. The risk that AI systems whose worst-case performance cannot be ensured or proven are implemented in high-stakes, safety-critical domains such as driving, medicine, and warfare, producing cascading errors and accidents that result in loss of life, economic damage, and social unrest. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1429 | 투명성과 해석 가능성 부족 Lack of transparency and interpretability 프론티어 AI가 해석하기 어렵고 투명성이 결여되며 학습 데이터에 대한 맥락적 이해가 명시적으로 내장되어 있지 않아, 미세조정이나 인간 피드백 기반 강화학습 없이는 소외 집단의 관점과 수행해야 할 한계를 포착하지 못하는 리스크. The risk that frontier AI is difficult to interpret and lacks transparency, with contextual understanding of the training data not explicitly embedded, so that without fine tuning or reinforcement learning with human feedback it fails to capture the perspectives of underrepresented groups or the limitations within which it is expected to perform. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1461 | 개발 문서화 미흡 Insufficient development documentation AI 시스템 개발 과정의 결정과 조치가 문서화되지 않아 개발 프로세스 최적화와 시스템의 감사가능성이 훼손되는 리스크. The risk that decisions and actions taken throughout the development of an AI system go undocumented, undermining optimization of the development process itself and the auditability of the AI system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1462 | 최종 사용자에게 부적절한 투명성 수준 Inappropriate degree of transparency to end users 최종 사용자에 대한 투명성이 설계에 적절히 통합되지 않아 올바른 운용이 저해되고 AI 애플리케이션의 오용이 유발되는 리스크. The risk that transparency to end users is not adequately integrated into the design of an AI system, preventing its proper operation and causing potential misuse of the AI application. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1463 | 하드웨어 연산·전력 요구사항 누락 Omitted compute and power hardware requirements AI 시스템의 개발과 운영에 필요한 상당한 연산·전력 요구가 하드웨어 선정에서 고려되지 않아 개발과 운영상 문제가 발생하는 리스크. The risk that the significant computational and power demands of AI system development and operation are not considered in hardware selection, creating issues in development and operation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1464 | 신뢰할 수 없는 데이터 소스 사용 Use of untrustworthy data sources 특히 제3자 데이터 소스를 활용할 때 신뢰할 수 없는 데이터 소스를 선택하여 데이터 품질 요구사항이 충족되지 못하는 리스크. The risk that choosing an untrustworthy data source, especially when third-party data sources are used to develop the AI system, leaves data quality requirements unfulfilled. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1465 | 데이터 이해 부족 Lack of data understanding 사용 데이터에 대한 이해 부족으로 데이터 결함이 간과되고 의도된 기능에 가장 적합한 AI 시스템의 개발이 저해되는 리스크. The risk that insufficient understanding of the data used for developing an AI system leaves data shortcomings unaddressed and hinders development of a system best suited to the intended functionality. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1466 | 잘못된 데이터 라벨 Incorrect data labels 데이터 레이블이 부정확하여 지도학습 AI 시스템이 실측 진실과 의도된 기능을 학습하지 못하는 리스크. The risk that incorrect data labels prevent a supervised learning AI system from learning the ground truth and therefore the intended functionality. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1472 | 과적합 및 과소적합 Over- and underfitting 모델이 훈련 데이터에 과도하게 또는 불충분하게 적응하여 운영 데이터에 직면했을 때 AI 시스템이 신뢰할 수 없게 동작하는 리스크. The risk that over- or underfitting, the excessive or insufficient adaption of a model to training data, causes an AI system to behave unreliably when confronted with operational data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1473 | 불투명성에 의한 결함 진단 저해 Opacity-impeded model debugging 블랙박스 모델 기반 AI 시스템의 설명가능성이 제한되어 개발자가 데이터나 모델 자체의 결함을 탐지하지 못하고 시스템의 성능과 안전 수준이 저하되는 리스크. The risk that the limited explainability of AI systems based on black-box models prevents developers from detecting shortcomings in the data or the model itself, decreasing the performance and safety levels of the AI system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1474 | 코너 케이스의 신뢰성 없음 Unreliability in corner cases AI 시스템이 희귀하거나 모호한 입력 데이터인 코너 케이스에 직면할 때 통제된 동작이 요구됨에도 신뢰할 수 없는 동작을 보이는 리스크. The risk that an AI system shows unreliable behavior when confronted with rare or ambiguous input data, also called corner cases, where controlled behavior is required. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1475 | 신뢰도 추정 기능 결여 Missing or faulty confidence estimation AI 시스템이 출력에 상응하는 신뢰도 수준을 제공하지 못하거나 잘못 제공하여 성능과 안전에 부정적 영향이 발생하는 리스크. The risk that an AI system fails to provide, or incorrectly provides, a level of confidence corresponding to its output, negatively impacting performance and safety. Source members (2)Source: min_cos=0.8331 · Mixed L3 RAI4-1475신뢰도 추정 기능 결여 RAI4-1717타당성·신뢰성 부족에 의한 부정확한 출력 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1476 | 운영 데이터 분포 편차 Operational data distribution deviation 테스트 세트가 근사한 분포와 실제 운영 데이터 분포 사이의 예기치 못한 편차로 배포된 AI 애플리케이션이 신뢰할 수 없게 동작하는 리스크. The risk that an unexpected deviation between the distribution approximated by the test set and the actual operational data causes a deployed AI application to behave unreliably. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1478 | 컨셉 드리프트 Concept drift 입력 변수와 모델 출력 간의 관계가 변화하는 개념 드리프트가 적절히 처리되지 않아 AI 시스템의 신뢰성이 저하되는 리스크. The risk that concept drift, a change in the relationship between input variables and model output, is not treated appropriately and reduces the reliability of AI systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1485 | AI 구성요소 상호작용의 원인 규명 곤란 Unclear attribution of harm from AI component interactions 서로 다른 AI 구성요소 간의 상호작용이 피해를 유발하지만 어떤 구성요소가 원인인지 특정하기 어려운 리스크. The risk that interactions between different AI components cause harm while it remains difficult to pinpoint which components are the cause. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1495 | 과도한 안전 튜닝 Overly restrictive safety tuning 과도한 안전 훈련이나 안전 튜닝이 AI 시스템의 성능을 저하시켜 지나치게 조심스러운 행동을 유발하고, 유해한 프롬프트와 부분적으로 유사한 완전히 안전한 프롬프트에 대해서도 응답을 거부하게 되는 리스크. The risk that excessive safety training or safety tuning impairs the performance of AI systems, leading to overly cautious behavior in which they refuse to answer entirely safe prompts that are partially similar to harmful ones. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1500 | 역량 식별·측정 곤란 Capability identification and measurement difficulty 범용 AI 시스템의 역량이 잠재적 위험의 분포가 넓고 이를 평가할 명확한 지표가 없으며 예측 불가능한 창발적 속성이 존재하여 고정 목적 AI에 비해 측정하기 어려운 리스크. The risk that the capabilities of general-purpose AI systems are difficult to measure compared with those of more limited and fixed-purpose AI systems, due in part to a broader distribution of potential risks, a lack of well-defined metrics to evaluate them, and risks from unpredictable or emergent model properties. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1502 | 내재 가치 측정 부정확 Inaccurate measurement of encoded values AI 시스템의 출력이 인간의 가치에 확고히 부합하는지 아니면 부분적으로만 상관된 모방인지 평가할 강건한 프레임워크가 부재하고, 모델이 학습한 가치 표상이 출력에 온전히 반영되지 않으며 훈련·배포 단계에 따라 어떻게 변하는지 알려지지 않아 내재 가치의 측정이 부정확해지는 리스크. The risk that, lacking robust frameworks for evaluating whether AI outputs robustly conform to human values rather than merely mimicking them, and with outputs imperfectly reflecting the learned value representations whose evolution across training and deployment stages is unknown, measurement of encoded values becomes inaccurate. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1504 | 인간 평가 한계 초과 출력 Outputs beyond human evaluability 인간 피드백을 이용한 평가로 AI 모델을 훈련할 때 평가자가 감지하기 어려운 오류를 포함한 출력을 정답과 유사하게 긍정 평가하여, 모델이 소프트웨어 취약점이 있는 코드나 정치적으로 편향된 정보처럼 미묘하게 잘못되거나 유해한 출력을 학습하고 극단적으로는 숨겨진 오류나 백도어를 포함한 출력을 생성하는 리스크. The risk that, when AI models are trained through evaluation with human feedback, human evaluators rate outputs containing hard-to-detect errors positively or similarly to correct ones, so the model learns to produce subtly incorrect or harmful outputs such as code with software vulnerabilities or politically biased information, and in extreme cases complicated outputs containing hidden errors or backdoors. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1543 | 아웃바운드 통신에 의한 기밀 유출·무단 행위 Data leakage and unauthorized actions from unintended outbound communication 네트워크 접근 권한을 폭넓게 가진 AI 시스템이 통신 채널 화이트리스트나 최소권한 원칙 없이 배포되어, 제공자·배포자·사용자가 의도하지 않은 아웃바운드 통신으로 기밀 데이터가 유출되거나 원치 않는 행위가 실행되는 리스크. The risk that an AI system with broad network access, deployed without communication whitelisting or least-privilege constraints, sends data outbound in ways no provider, deployer, or user intended, leaking confidential data or performing unwanted actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1551 | 센서 드리프트에 의한 배포 시스템 성능 저하 Deployed-system degradation from sensor and distribution drift 물리 센서와 데이터 소스에 의존하는 배포된 AI 시스템에서 하드웨어 드리프트에 따른 데이터 분포 변화가 발생하여 시스템의 견고성과 성능이 저하되는 리스크. The risk that hardware drift in the physical sensors and data sources of deployed AI systems causes data distribution drift that degrades system robustness and performance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1633 | 인프라 제어 오판단에 의한 필수 서비스 붕괴 Essential-service collapse from erroneous AI infrastructure control 전력망·수처리·통신·교통 조정 시스템에 배포된 범용 AI가 운영 데이터를 오해석하거나 연쇄 장애를 예측하지 못한 제어 결정을 내려 상호연결된 인프라가 불안정해지고 광범위한 정전, 수질 오염, 통신 두절 등 필수 서비스가 붕괴되는 리스크. The risk that general-purpose AI deployed in power grid, water treatment, telecommunications, or transportation systems misinterprets operational data, fails to anticipate cascading failure modes, or makes control decisions that destabilize interconnected infrastructure, causing widespread blackouts, contaminated water supplies, communications breakdowns, and collapse of essential services. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1703 | AI 시스템 요구사항·목적 명세 오류 Misspecification of AI system requirements and purpose AI 시스템의 요구사항과 목적이 잘못 이해되거나 잘못 도출되어 부정확한 요구사항 도출, 차선의 모델링·하이퍼파라미터 선택, 변화하는 도메인 요구에 대한 적응 실패가 발생하는 리스크. The risk that the requirements and purpose of an AI system are misunderstood or mis-elicited, resulting in incorrect requirement elicitation, suboptimal modelling or hyperparameter choices, and failure to adapt to changing domain requirements. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1705 | 모니터링 부족에 의한 미탐지 안전·프라이버시 위반 Undetected safety and privacy violations from insufficient monitoring 배포된 AI에 대한 모니터링과 해석 가능성이 부족하여 블랙박스 불투명성이 인간의 주체성을 축소하고 윤리·안전 원칙 위반과 프라이버시 침해가 탐지되지 않은 채 남는 리스크. The risk that insufficient monitoring and interpretability of deployed AI leaves black-box opacity diminishing human agency and allows ethical or safety-principle violations and privacy violations to go undetected. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1719 | 보안·복원력 부족에 의한 공격 취약과 복구 실패 Attack susceptibility and recovery failure from lack of security and resilience AI 시스템이 적대적 공격, 데이터 오염, 정보 유출에 취약하거나 악영향을 견디고 회복하지 못하는 리스크. The risk that an AI system is susceptible to adversarial attacks, data poisoning, or exfiltration, or is unable to withstand and recover from adverse events. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1721 | 설명 가능성 및 해석 가능성 부족 Lack of explainability and interpretability AI 시스템 작동의 기저 메커니즘을 표현하거나 출력의 의미를 맥락에 맞게 전달하지 못하는 리스크. The risk of inability to represent the mechanisms underlying an AI system's operation or to convey the meaning of its outputs in context. | ① Description ② L3 mapping ③ Duplicate |
상호작용 안전성 · Interaction Safety · 219 cards
RAI3-G-INT-01 폭력 Violence10 cards
타인·집단·동물에 대한 물리적·정신적 해를 가하거나, 그 위협·조장·미화를 포함하는 콘텐츠
| ID | Card | Human audit |
|---|---|---|
| RAI4-0426 | 혐오발언 생성 Hate speech generation 생성 시스템이 특정 정체성 집단을 표적으로 하는 혐오적·모욕적 콘텐츠를 산출하는 리스크. The risk that generative systems produce hateful or abusive content targeting identity groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0674 | 문화적 박탈 Cultural dispossession 말하기 방식, 유머 표현, 문화 정체성을 구성하는 소리와 목소리 등 문화적 재화와 가치가 의도적 또는 비의도적으로 소거되거나 다른 문화에서 부적절하게 재사용되는 리스크 The risk of intentional or unintentional erasure of cultural goods and values, such as ways of speaking, expressing humour, or sounds and voices that contribute to a cultural identity, or their inappropriate re-use in other cultures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0678 | 신체적 위해를 유발하는 출력 Model outputs leading to physical harm 모델이 명백히 폭력적이거나 은밀하게 위험하거나 그 밖에 간접적으로 안전하지 않은 언어를 생성하여 신체적 위해로 이어지는 리스크 The risk that a model generates overtly violent, covertly dangerous, or otherwise indirectly unsafe language that leads to physical harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0705 | 폭력 조장 위험 콘텐츠 Violence-inciting dangerous content 폭력적·선동적·급진화·위협적 콘텐츠의 제작과 접근이 용이해지고 자해나 불법 활동의 수행을 권장하는 콘텐츠가 산출되며, 증오·비하·고정관념 콘텐츠에 대한 대중의 노출을 통제하기 어려워지는 리스크 The risk of eased production of and access to violent, inciting, radicalizing, or threatening content and recommendations to carry out self-harm or conduct illegal activities, including difficulty controlling public exposure to hateful, disparaging, or stereotyping content. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0727 | 개인·집단·조직에 대한 언어적 공격 Verbal attacks on individuals, groups, or organizations 챗봇이 개인, 집단 또는 조직을 언어적으로 공격하거나 훼손하는 리스크 The risk that a chatbot verbally attacks or undermines an individual, group, or organization. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1049 | 조직 재정·평판 손상 Organizational financial and reputational damage ML 시스템을 구축하거나 사용하는 조직의 재정적 및/또는 평판 손상 위험. The risk of financial and/or reputational damage to the organization building or using the ML system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1070 | 유독한 언어 Toxic language LM이 욕설, 정체성 공격, 모욕, 위협, 성적으로 노골적인 내용, 비하 표현, 폭력 선동 등 증오 표현이나 유독한 언어를 예측·생성하는 리스크. The risk that LMs predict hate speech or other toxic language, including profanities, identity attacks, insults, threats, sexually explicit content, demeaning language, or language inciting violence targeted at a person or group because of innate characteristics. Source members (4)Source: min_cos=0.7697 · Mixed L3 RAI4-0726학습 데이터 내 독성 언어 RAI4-0729정체성 공격형 유해 언어 RAI4-1052증오심 표현 및 공격적인 언어 RAI4-1070유독한 언어 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1177 | 공격적 콘텐츠 생성 Offensive content generation LLM이 위협과 모욕, 경멸, 욕설, 빈정거림, 무례함 같은 공격적 콘텐츠나 행위를 식별하고 반대하지 못한 채 그러한 콘텐츠를 생성하는 리스크. The risk that LLMs fail to identify and oppose offensive content or actions involving threat, insult, scorn, profanity, sarcasm, and impoliteness, and generate such content instead. Source members (4)Source: min_cos=0.7998 RAI4-0636LLM 악용에 의한 독성 콘텐츠 생성 RAI4-0690LLM의 명시적·암묵적 독성 콘텐츠 생성 RAI4-1177공격적 콘텐츠 생성 RAI4-1188폭력적 콘텐츠 생성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1308 | 독성 텍스트 생성 Toxic text generation LLM이 프롬프트를 받았을 때 혐오 발언과 모욕적 언어, 폭력적 발언, 욕설을 아우르는 독성 텍스트를 생성하는 리스크. The risk that an LLM generates toxic text when prompted, toxicity being an umbrella term encompassing hate speech, abusive language, violent speech, and profane language. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1447 | 폭력/무력 충돌 Violence/armed conflict 기술 시스템을 사용하거나 오용하여 사이버 공격, 보안 침해, 치명적 생화학 무기 개발을 선동·촉진·수행함으로써 폭력과 무력 충돌이 초래되는 리스크. The risk that use or misuse of a technology system to incite, facilitate, or conduct cyberattacks, security breaches, and lethal biological and chemical weapons development results in violence and armed conflict. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-02 성적 콘텐츠 Sexual17 cards
성적 행위·성범죄·성착취·아동 성적 콘텐츠·성상품화 등 성 관련 유해 콘텐츠
| ID | Card | Human audit |
|---|---|---|
| RAI4-0627 | 허위 콘텐츠에 의한 개인 표적 피해 Targeted harm to individuals via fake content 악의적 행위자가 범용 AI로 허위 콘텐츠를 생성하여 사기, 갈취, 심리적 조작, 비동의 성적 이미지와 아동 성착취물 제작, 개인·조직에 대한 표적 방해에 사용함으로써 개인에게 표적 피해를 입히는 리스크 The risk that malicious actors use general-purpose AI to generate fake content that harms individuals in a targeted way, including scams, extortion, psychological manipulation, non-consensual intimate imagery and child sexual abuse material, and targeted sabotage of individuals and organisations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0641 | 비합의 성적 대상화 Non-consensual sexualization 기술이나 애플리케이션을 이용해 개인 또는 집단을 동의 없이 성적으로 대상화하는 리스크 The risk of the non-consensual sexualisation of an individual or group using a technology or application. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0642 | 증오·모욕·외설 콘텐츠 생성 Intentional generation of hateful and obscene content 생성형 AI 모델이 증오·모욕·모독(HAP) 또는 외설적 콘텐츠를 생성하는 데 의도적으로 사용되는 리스크 The risk that generative AI models are used intentionally to generate hateful, abusive, and profane (HAP) or obscene content. Source members (2)Source: min_cos=0.8589 · Mixed L3 RAI4-0642증오·모욕·외설 콘텐츠 생성 RAI4-0689증오·모욕·외설 출력 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0670 | 집단 왜곡 표상과 독성 콘텐츠 생성 Group misrepresentation and toxic content generation AI 시스템이 특정 집단을 과소·과대 표현하거나 허위로 표현하고 독성·모욕·학대·증오 콘텐츠를 생성하는 리스크 The risk that AI systems under-, over-, or misrepresent certain groups or generate toxic, offensive, abusive, or hateful content. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0684 | 성범죄 조장 콘텐츠 생성 Generation of content enabling sex-related crimes AI가 성매매 인신매매, 성폭력, 성희롱, 비합의 친밀 콘텐츠 유포, 수간 등 성 관련 범죄의 실행을 가능하게 하거나 조장 또는 승인하는 응답을 산출하는 리스크 The risk that an AI system produces responses that enable, encourage, or endorse the commission of sex-related crimes such as sex trafficking, sexual assault, sexual harassment, nonconsensual sharing of sexually intimate content, and bestiality. Source members (6)Source: min_cos=0.7477 · Mixed L3 RAI4-0665AI 기반 아동 성착취물 생성 RAI4-0673아동 성착취 조장 콘텐츠 생성 RAI4-0682비폭력 범죄 조장 콘텐츠 생성 RAI4-0684성범죄 조장 콘텐츠 생성 RAI4-0685음란물 및 성적 대화 콘텐츠 생성 RAI4-0691폭력 범죄 조장 콘텐츠 생성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0688 | 커뮤니티 기준 위반 독성 콘텐츠 생성 Generation of community-standard-violating toxic content 유혈, 아동 성적 묘사, 욕설, 정체성 공격 등 커뮤니티 기준을 위반하는 콘텐츠가 생성되어 특정 집단에 피해를 주거나 그들에 대한 증오와 폭력을 선동하는 리스크 The risk of generating content that violates community standards, including harming or inciting hatred or violence against groups, such as gore, sexual content of children, profanities, and identity attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0714 | 아동·청소년 유해 콘텐츠 제공 Provision of content harmful to children and youth LLM이 아동과 청소년에게 유해한 콘텐츠를 담은 답변을 산출하도록 유도될 수 있는 리스크 The risk that LLMs are leveraged to solicit answers that contain content harmful to children and youth. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0718 | 외설·비하·학대 이미지 생성 및 접근 용이화 Eased production of and access to abusive imagery 해를 끼칠 수 있는 외설적·굴욕적·모욕적 이미지, 특히 합성 아동 성적 학대 자료(CSAM)와 성인의 비동의 친밀 이미지(NCII)의 제작과 접근이 용이해지는 리스크 The risk that production of and access to obscene, degrading, or abusive imagery which can cause harm is eased, including synthetic child sexual abuse material (CSAM) and nonconsensual intimate images (NCII) of adults. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0728 | 독성 콘텐츠 생성 Generation of toxic content 모델이 무례하고 모욕적이며 심지어 불법적인 정보를 담은 독성 콘텐츠를 생성하는 리스크 The risk that a model generates toxic content containing rude, disrespectful, and even illegal information. Source members (7)Source: min_cos=0.7575 · Mixed L3 RAI4-0710유해·차별적 콘텐츠 생성 RAI4-0712비윤리적·유해 콘텐츠 및 고위험 조언 생성 RAI4-0728독성 콘텐츠 생성 RAI4-0936독성 및 악의적인 콘텐츠 RAI4-1165모욕적 콘텐츠 생성 RAI4-1220유해·부적절 콘텐츠 생성 RAI4-1349의도치 않은 유해 콘텐츠 생성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0828 | 참여 극대화 콘텐츠에 의한 자율성 훼손 Erosion of user autonomy from engagement-maximizing content 생성형 AI가 진실성과 무관하게 클릭베이트 헤드라인과 대량의 기사·웹페이지를 생성하여 검색 노출과 클릭을 극대화함으로써, 이용자의 탐색 행동을 조작하고 이용 경험과 소비자 자율성을 훼손하는 리스크 The risk that generative AI produces clickbait headlines and mass articles regardless of veracity to maximize search visibility and clicks, manipulating how users navigate the internet and applications, degrading their experience, and undermining consumer autonomy. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1008 | 기술을 이용한 폭력 Technology-facilitated violence 알고리즘 기능이 시스템을 괴롭힘과 폭력에 이용할 수 있게 하여 생성 AI의 비합의 성적 이미지 생성, 신상털기, 트롤링, 사이버스토킹·괴롭힘, 감시·통제 등 온라인 폭력이 발생하는 리스크. The risk that algorithmic features enable use of a system for harassment and violence, including non-consensual sexual imagery in generative AI, doxxing, trolling, cyberstalking, cyberbullying, monitoring and control, and online harassment and intimidation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1136 | 대규모 유해 콘텐츠 생성 Scaled harmful content generation 적절한 안전·보안 장치 없이 고급 AI 비서가 위협 행위자로 하여금 아동 성학대 자료, 사기, 허위정보 같은 유해 콘텐츠를 더 빠르고 정확하며 저렴하고 개인화된 형태로 더 넓은 범위에 생성하게 하는 리스크. The risk that, without proper safety and security mechanisms, advanced AI assistants allow threat actors to create harmful content such as child sexual abuse material, fraud, and disinformation more quickly, accurately, cheaply, and with greater personalization and reach. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1137 | 합의되지 않은 콘텐츠 대규모 생성 Non-consensual content generation at scale 생성 AI가 나체, 혐오, 폭력 묘사와 개인의 초상을 이용한 합의되지 않은 콘텐츠를 만들고, 비서의 도구 사용과 계획 능력이 개인 표적화와 착취·괴롭힘·협박을 자동화하고 대규모로 확산시키는 리스크. The risk that generative AI produces non-consensual content depicting nudity, hate, or violence and using individuals' likenesses, while assistants' tool-use and planning capabilities automate targeting of individuals for exploitation, harassment, and blackmail at scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1204 | 비합의 성적 딥페이크 Non-consensual sexual deepfakes 생성 AI가 합의되지 않은 성적 이미지와 영상의 딥페이크를 만들어 타인을 해치거나 굴욕감을 주거나 성적으로 대상화하는 데 악의적으로 사용되는 리스크. The risk that generative AI is maliciously used to harm, humiliate, or sexualize another person by generating deepfakes of nonconsensual sexual imagery or videos. Source members (3)Source: min_cos=0.7972 · Mixed L3 RAI4-0638동의 없는 딥페이크 인물 모사 RAI4-1204비합의 성적 딥페이크 RAI4-1357성적 노골 콘텐츠 오용 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1542 | 외부 도구 연동을 통한 유해 콘텐츠 유입 Ingestion of harmful content through external tool integration 외부 도구·플러그인과의 통합과 상호 연결이 확대됨에 따라 악의적 외부 입력에 노출되어 유해 콘텐츠가 시스템에 유입되는 리스크. The risk that growing integration and interconnectivity with external tools and plugins exposes a system to malicious external inputs that introduce harmful content. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1555 | 취약점 기반 맞춤형 괴롭힘 콘텐츠 생성 Vulnerability-targeted personalized harassment content generation GPAI가 표적 개인의 약점에 맞춘 콘텐츠를 자동 생성하는 데 오용되어 괴롭힘·갈취·협박의 효율성과 성공률이 높아지는 리스크. The risk that GPAI is misused to automatically generate content personalized to targets' weak spots, making harassment, extortion, and intimidation more efficient and more likely to succeed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1582 | 비동의 성적 이미지 생성 Non-consensual intimate imagery generation 성인의 실제 외모를 이용한 노골적 성적 자료가 본인의 동의 없이 생성되는 리스크. The risk that sexually explicit material is created using an adult person's likeness without their consent. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-03 자해 Self-harm5 cards
자살·자해·위험 약물 남용·극단적 다이어트 등 개인의 신체·정신 안전을 직접 위협하는 콘텐츠
| ID | Card | Human audit |
|---|---|---|
| RAI4-0023 | 에이전트의 자해 조장 Self-harm facilitation by agents 에이전트가 자해 또는 대인 위해의 가능성을 높이는 행동 지향적 지원·계획·자원을 제공하는 리스크. The risk that an agent provides action-oriented support, planning, or resources that increase the likelihood of self-harm or interpersonal harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0157 | 챗봇을 통한 자해 격려 Self-harm encouragement by chatbot 챗봇이 자해 사고나 행동을 강화하거나 정상화하거나 중단시키지 못하는 리스크. The risk that a chatbot reinforces, normalizes, or fails to interrupt self-harm ideation or behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0687 | 자살·자해 조장 콘텐츠 생성 Generation of content encouraging suicide and self-harm AI가 자살, 자해, 섭식장애 등 의도적 자해 행위를 가능하게 하거나 조장 또는 승인하는 응답을 산출하는 리스크 The risk that an AI system produces responses that enable, encourage, or endorse acts of intentional self-harm such as suicide, self-injury, and disordered eating. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0833 | 정신건강 유해 콘텐츠 Mental health harmful content 모델이 자살을 조장하거나 공황·불안을 유발하는 등 정신건강에 위험한 응답을 생성하여 이용자의 정신건강에 부정적 영향을 미치는 리스크 The risk that a model generates risky responses about mental health, such as content that encourages suicide or causes panic or anxiety, negatively affecting users' mental health. Source members (2)Source: min_cos=0.8349 RAI4-0833정신건강 유해 콘텐츠 RAI4-0839신체 건강 위해 콘텐츠 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0861 | 자해 촉진 Self-harm facilitation 기술 시스템 사용의 직접적 또는 간접적 결과로 이용자가 자신의 신체에 고의적 손상을 가하게 되는 리스크 The risk that a person deliberately damages their own body as a direct or indirect result of using a technology system. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-04 혐오·차별 Hate and Unfairness9 cards
인종, 지역, 국적, 민족, 가정형태, 공공, 성, 세대/나이, 신체 조건, 유명인, 출산 및 혼인 여부, 장애/병력, 재난 및 범죄 피해, 종교/신념, 직업/학력/사회적 지위, 취미 및 욕설/비속어등 + 종교
| ID | Card | Human audit |
|---|---|---|
| RAI4-0411 | 언어별 피해 과소탐지 Language-specific harm under-detection 평가 범위가 취약한 언어와 방언에서 유해·모욕적·비안전 콘텐츠가 덜 안정적으로 탐지되는 리스크. The risk that toxic, disrespectful, or unsafe content is less reliably detected in languages and dialects with weaker evaluation coverage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0683 | 출력 단계 집단 표상 편향 Group misrepresentation bias in model outputs 생성된 콘텐츠가 특정 집단이나 개인을 불공정하게 표상하는 리스크 The risk that generated content unfairly represents certain groups or individuals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0706 | 보호 속성 기반 부당 대우 Protected-attribute unfair treatment 인종, 민족, 연령, 성별, 성적 지향, 종교, 출신 국가, 혼인 여부, 장애, 언어 등 보호 속성을 근거로 개인이 불공정하거나 부적절한 대우 또는 자의적 차별을 받는 리스크 The risk of unfair or inadequate treatment or arbitrary distinction based on a person's race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other protected groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0720 | 편견과 과소표현으로 인한 위험 Risks from bias and underrepresentation 영어권·서구 편중 학습 데이터로 인해 인종, 성별, 문화, 연령, 장애 전반에 편향된 출력이 산출되어 의료, 채용, 대출 등 고위험 영역에서 과소대표 사용자에게 피해를 주는 리스크 Training corpora that overrepresent English-speaking and Western populations produce outputs biased across race, gender, culture, age, and disability, harming underrepresented users in high-stakes domains such as healthcare, recruitment, and lending. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0937 | 사회 집단에 대한 부당한 부정적 편견 Unfairly negative bias against social groups 모델이 일방적이거나 부정확한 정보에 기반해 성별·인종·종교 등에 대한 부정적 고정관념과 결부된, 사회 집단이나 개인에 대한 부당하게 부정적인 태도를 드러내는 리스크. The risk that a system exhibits an unfairly negative attitude towards a social group or individuals based on one-sided or inaccurate information, typically tied to widely disseminated negative stereotypes regarding gender, race, or religion. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1068 | 고정관념 인코딩에 의한 차별적 대우 Discriminatory treatment from encoded stereotypes 차별적 언어와 사회적 고정관념을 인코딩한 언어모델이 성별·종교·성적지향·장애·연령 등 민감한 속성에 따른 차별적 대우와 자원 접근 격차 등 다양한 피해를 야기하는 리스크. The risk that language models encoding discriminatory language or social stereotypes cause harms including differential treatment or access to resources based on sensitive traits such as sex, religion, gender, sexual orientation, ability, and age. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1166 | 불공정·차별적 산출 Unfair and discriminatory output 모델이 인종과 성별, 종교, 외모 등에 근거한 사회적 편향을 비롯한 불공정하고 차별적인 산출을 생성하여 특정 집단을 불편하게 하고 사회의 안정과 평화를 훼손하는 리스크. The risk that a model produces unfair and discriminatory data, such as social bias based on race, gender, religion, or appearance, which may discomfort certain groups and undermine social stability and peace. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1168 | 민감 주제 편향 콘텐츠 Biased content on sensitive topics 정치를 비롯한 민감하고 논쟁적인 주제에서 언어모델이 특정 정치적 입장을 지지하는 편향되고 오도하며 부정확한 콘텐츠를 생성하여 다른 관점을 차별하거나 배제하는 리스크. The risk that on some sensitive and controversial topics, especially politics, language models generate biased, misleading, and inaccurate content that supports a specific position and leads to discrimination or exclusion of other viewpoints. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1174 | 안전하지 않은 지시 주제 Unsafe instruction topic 입력 지시 자체가 부적절하거나 불합리한 주제를 참조할 때 모델이 그 지시를 따라 광신과 인종주의 같은 안전하지 않은 콘텐츠를 생성하여 사회에 부정적 영향을 줄 수 있는 리스크. The risk that when the input instructions themselves refer to inappropriate or unreasonable topics, the model follows them and produces unsafe content such as fanaticism or racism with possible negative impact on society. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-05 정치적 중립성 Political Neutrality6 cards
정치적 극단주의 선동, 정치인 비방 및 명예훼손, 선거 절차 방해와 선동, 종교 전복 및 관련 사기 수법, 정치적 동기에 의한 폭력의 정당화,
| ID | Card | Human audit |
|---|---|---|
| RAI4-0403 | 도덕적 프레이밍 편향 Moral framing bias 프롬프트·응답·평가의 구성 방식이 이용자를 논쟁적 사안에 대한 특정한 도덕적 해석으로 유도하는 리스크. The risk that the framing of prompts, responses, or evaluations channels users toward a particular moral interpretation of a contested issue. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0422 | 도덕적 세계관 배제 Moral-worldview exclusion 종교적·영적·토착적·비세속적 도덕 세계관이 모델 행동과 평가 기준에서 배제되는 리스크. The risk that religious, spiritual, indigenous, or non-secular moral worldviews are excluded from model behavior and evaluation standards. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0529 | 정치 불안정 유발 오용 Politically destabilizing misuse 기술 시스템의 사용 또는 오용이 직간접적으로 정치적 불안을 야기하는 리스크 The risk that the use or misuse of a technology system directly or indirectly causes political unrest. Source members (2)Source: min_cos=0.8352 RAI4-0529정치 불안정 유발 오용 RAI4-1450구조적 정치 불안정화 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0939 | 모델의 극단적·편향적 견해 표출 Expression of extremist and politically biased views 대형 모델이 정치적 주제에서 부적절하거나 극단주의적인 견해를 표출하고, 중립을 표방하면서도 특정 정치 성향의 편향을 드러내는 리스크. The risk that large models express inappropriate or extremist views on political topics and, while claiming neutrality, exhibit notable political biases across policy domains. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1159 | 정치적 영향력 전략 역량 Political influence strategy capability 모델이 행위자가 정치적 영향력을 획득하고 행사하는 데 필요한 사회적 모델링과 계획을, 다수 행위자와 풍부한 사회적 맥락이 있는 시나리오에서까지 수행하는 리스크. The risk that a model performs the social modelling and planning necessary for an actor to gain and exercise political influence, not just at a micro level but in scenarios with multiple actors and rich social context. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1451 | 정치적 조작 Political manipulation 개인 데이터를 사용하거나 오용하여 마이크로 광고나 딥페이크·합성 미디어를 통해 개인의 관심사·성격·취약성을 표적으로 맞춤형 정치 메시지를 전달하는 리스크. The risk that personal data is used or misused to target individuals' interests, personalities, and vulnerabilities with tailored political messages via micro-advertising or deepfakes and synthetic media. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-06 개인정보 Privacy39 cards
개인식별정보 노출, 의료/금융/위치/생체 정보 노출, 통신내용 침해, 개인사 노출, 개인 신상 정보 수집 방법, 프라이버시 침해 기술
| ID | Card | Human audit |
|---|---|---|
| RAI4-0433 | 생체정보 기반 프라이버시 침해 Biometric privacy intrusion 얼굴·음성·보행 등 생체인식 AI 시스템이 침해적인 식별과 추적을 가능하게 하는 리스크. The risk that face, voice, gait, or other biometric AI systems enable intrusive identification and tracking. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0631 | 신원 사칭 및 도용 Impersonation and identity theft 제3자가 개인, 집단, 조직의 신원을 도용하여 사기, 조롱, 그 밖의 가해를 함으로써 당사자나 다른 당사자에게 피해가 발생하는 리스크 The risk that a third party steals the identity of an individual, group, or organisation in order to defraud, mock, or otherwise harm them or another party. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0652 | 개인화 정보 악용 표적 사기 Targeted fraud exploiting personalized information 생성 모델이 개인화된 정보를 이용해 개별 사용자를 효율적으로 표적화하는 데 오용되어, 피해자의 신뢰를 악용한 설득력 높은 자동화 사기로 민감 정보가 탈취되고 기만의 성공 가능성이 높아지는 리스크 The risk that generative models are misused to target individual users more efficiently using personalized information, producing highly convincing automated fraudulent schemes that exploit victims' trust and extract sensitive data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0655 | 초상 도용 Appropriation of personal likeness 개인의 초상이나 그 밖의 식별 가능한 특징이 사용되거나 변형되는 리스크 The risk that a person's likeness or other identifying features are used or altered. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0734 | 기밀정보 무단 공유에 따른 사업 손실 Business loss from unauthorised sharing of confidential information 기업 전략, 재무 계획 등 민감·기밀 정보와 문서가 제3자와 무단으로 공유되어 시장 지위나 수익을 상실하는 리스크 The risk that sensitive, confidential information and documents such as corporate strategy and financial plans are shared with third parties without authorisation, risking loss of market position or revenue. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0735 | 개인 데이터의 공개 및 부적절한 공유 Disclosure and improper sharing of personal data 개인의 데이터가 공개되거나 부적절하게 공유되고, AI가 원시 데이터에 명시적으로 담기지 않은 추가 정보를 추론하거나 모델 학습을 위해 개인 데이터를 공유함으로써 새로운 유형의 공개 위험이 생기고 악화되는 리스크 The risk that data of individuals is revealed and improperly shared, with AI creating new types of disclosure risk by inferring additional information beyond what is explicitly captured in the raw data and exacerbating disclosure risks through sharing personal data to train models. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0739 | 민감한 정보 노출 Sensitive-information exposure 은폐하도록 사회화된 매우 사적인 민감 정보가 드러나고, 생성 기술이 검열·삭제된 콘텐츠를 재구성하거나 추론된 민감 데이터·선호·의도를 노출함으로써 새로운 유형의 노출 위험이 발생하는 리스크 The risk that sensitive private information people view as deeply primordial and have been socialized into concealing is revealed, with AI creating new types of exposure risk through generative techniques that reconstruct censored or redacted content and through exposing inferred sensitive data, preferences, and intentions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0741 | 개인정보 보호조치 미흡 Inadequate personal-data protection 결함 있는 데이터 저장·처리 관행이 수집된 개인정보를 유출과 부적절한 접근으로부터 보호하지 못하는 리스크 The risk that faulty data storage and handling practices fail to protect collected personal data from leaks and improper access. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0751 | 민감 개인정보 출력 노출 Exposure of sensitive personal information in outputs AI가 자택 주소, 로그인 자격증명, 계좌·카드 정보 등 비공개 민감 개인정보를 출력하여 개인의 물리적·디지털·재정적 안전을 위협하는 리스크 The risk that an AI system outputs sensitive, non-public personal information such as home addresses, login credentials, or financial account details, undermining a person's physical, digital, or financial security. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0760 | 프롬프트 프라이밍에 의한 개인정보 유도 출력 Personal-data elicitation through prompt priming 생성 모델이 입력과 유사한 출력을 산출하는 성질을 이용해 프롬프트에 개인정보를 포함시킴으로써, 학습 데이터에 포함되었던 개인정보가 출력으로 드러나는 리스크 The risk that, because generative models produce output like the input provided, priming a prompt with personal information elicits similar personal data that was included in the model's training. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0761 | 익명화된 데이터의 재식별 Re-identification of anonymized data 데이터에서 개인식별정보(PII)와 민감 개인정보(SPI)를 제거하더라도 데이터에 남아 있는 다른 특성과의 상관관계로 인해 개인이 재식별되는 리스크 The risk that, even with the removal of personally identifiable information (PII) and sensitive personal information (SPI) from data, persons can be identified due to correlations to other features available in the data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0764 | 개인정보의 무단 사용과 오용 Unauthorized use and misuse of personal information AI 시스템이 민감한 개인정보를 수집·보관·활용하는 과정에서 투명성과 보호장치가 미흡하여 개인 데이터가 무단으로 사용되거나 오용되는 리스크 The risk that AI systems acquire, store, and use sensitive personal information without adequate transparency or safeguards, resulting in unauthorized exploitation or misuse of personal data. Source members (12)Source: min_cos=0.6707 · Mixed L3 RAI4-0764개인정보의 무단 사용과 오용 RAI4-0765동의 없는 개인정보 2차 이용 RAI4-0988AI 시스템 해킹·무단 조작에 의한 오용 RAI4-1013남용 및 오용 RAI4-1038AI의 오용 RAI4-1073개인정보를 정확하게 추론한 프라이버시 침해 RAI4-1099무단 개인데이터 수집 RAI4-1335불법 데이터 수집·사용 RAI4-1407악의적 프라이버시 침탈 RAI4-1624개인정보 노출과 책임 리스크 RAI4-1625독점 데이터의 무단 접근·복제·공개 RAI4-1714AI 시스템에 의한 개인 프라이버시 침해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0772 | 개인정보 연계에 의한 프라이버시 침해 Privacy violation through personal-information association LLM이 특정 개인과 관련된 여러 개인식별정보를 연계·기억하여, 한 정보를 단서로 제시하는 프롬프트만으로 이메일 등 다른 개인정보가 출력되는 리스크 The risk that an LLM associates and retains multiple pieces of personally identifiable information about a person, so that a prompt referencing one item elicits disclosure of another, such as an email address. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0777 | 개인정보·민감 데이터 유출과 무단 이용 Leakage and unauthorized use of personal data 생체인식·건강·위치 등 개인식별정보나 민감 데이터가 유출되거나 무단으로 이용·공개되거나 익명성이 해제되어 영향이 발생하는 리스크 The risk of impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0781 | 추론된 개인정보 기반 의사결정 피해 Harm from decisions based on inferred private data 범용 AI가 이용자가 제공한 맥락 입력으로부터 민감한 개인정보를 높은 정확도로 추론하고 이를 근거로 의사결정에 활용함으로써, 정보 누출·불공정 대우·행동 조작이 발생하는 리스크 The risk that general-purpose AI infers sensitive information about users from contextual input and acts on those inferences, leaking private information, causing unfair treatment, or enabling manipulation of user behaviour. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0837 | 허위정보 및 개인정보 침해 Misinformation and privacy violations 신뢰성이 낮은 범용 모델이 허위·오도 정보를 유포하거나 핵심 정보를 누락하고, 사실 정보라도 프라이버시권을 침해하는 방식으로 전달하는 리스크 Unreliable general-purpose models disseminate false or misleading information, omit critical information, or reveal true information in ways that violate privacy rights. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0914 | 생성 출력의 개인정보 유출 Privacy leakage in generated output 생성된 콘텐츠에 민감한 개인정보가 포함되어 유출되는 리스크. The risk that generated content includes sensitive personal information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0922 | 학습 코퍼스 내 개인정보 혼입 Private data contamination of training corpora 웹 수집 데이터와 인간-기계 대화 데이터의 통합 과정에서 이름·이메일·주소 등 개인식별정보(PII)가 학습 코퍼스에 혼입되어 오용되는 리스크. The risk that personally identifiable information such as names, emails, and addresses is mixed into training corpora through web-collected data and human-machine conversations, enabling misuse. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0933 | 프라이버시 및 데이터보호 규정 위반 Privacy and data protection regulation violations AI 시스템이 대규모 개인정보 수집·저장, 데이터 유출, 무단 웹 스크레이핑 등을 통해 프라이버시를 침해하고 GDPR 등 데이터보호 규정을 위반하는 리스크. The risk that AI systems violate privacy and data protection regulations such as GDPR through mass collection and storage of personal data, data breaches, and unauthorized web scraping. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0976 | 얼굴인식 기술에 의한 프라이버시 침해 Privacy harms from face recognition technologies 얼굴 인식 기술과 그 유사 기술이 저장 데이터의 보관 기간·소유·법적 소환 가능성 등에 걸쳐 심각한 프라이버시 침해를 초래하는 리스크. The risk that face recognition technologies and their ilk pose significant privacy harms, raising unresolved questions about what data is stored, for how long, who owns it, and whether it can be subpoenaed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1047 | ML 시스템의 개인정보 유출 피해 Harm from privacy leakage in ML systems ML 시스템을 통한 개인정보 유출로 인한 손실 또는 피해 위험. The risk of loss or harm from leakage of personal information via the ML system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1055 | 민감한 정보 유출로 인한 개인정보 침해 Compromising privacy by leaking sensitive information 학습 데이터에 개인정보가 존재할 경우 LM이 이를 기억했다가 유출하여 프라이버시가 침해되는 리스크. The risk that an LM remembers and leaks private data present in its training data, causing privacy violations. Source members (6)Source: min_cos=0.7083 · Mixed L3 RAI4-0923암기된 학습 데이터의 유출 RAI4-1055민감한 정보 유출로 인한 개인정보 침해 RAI4-1072개인정보 유출로 인한 사생활 침해 RAI4-1074민감한 정보의 유출이나 정확한 추론으로 인한 위험 RAI4-1144가중치 기억에 의한 개인정보 유출 RAI4-1591프라이버시 공격에 의한 훈련 데이터 민감정보 노출 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1062 | 사용자 신뢰 악용을 통한 사적 정보 유도 Exploiting user trust to elicit private information 대화 중 사용자가 의견이나 감정 등 평소 얻기 어려운 사적 정보를 털어놓게 되고(인간처럼 보이는 챗봇일수록 더 많이 공개), 그 정보가 프라이버시권을 침해하거나 중독성 애플리케이션 추천 등으로 사용자에게 해를 끼치는 후속 응용에 이용되는 리스크. The risk that users reveal private information in conversation, such as opinions or emotions, that would otherwise be difficult to access—disclosing more to human-like chatbots—enabling downstream applications that violate privacy rights or cause harm, for example through more effective recommendation of addictive applications. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1089 | 개인 초상·신원 무단 사용 Non-consensual use of personal identity or likeness 상업적 목적 등 승인되지 않은 목적을 위해 개인의 신원이나 초상을 동의 없이 사용하는 리스크. The risk of non-consensual use of a person's identity or likeness for unauthorised purposes such as commercial purposes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1139 | 비서 유도 개인정보 노출 Assistant-induced privacy disclosure 비서가 사용자로 하여금 자신이나 타인의 개인정보를 공개하도록 유도해 신원 도용과 낙인·차별을 초래하고, 국가 소유 비서가 조작이나 기만으로 감시 목적의 사적 정보를 추출하는 리스크. The risk that assistants influence users to disclose personal information or private information pertaining to others, resulting in identity theft, stigmatisation and discrimination, and that state-owned assistants employ manipulation or deception to extract private information for surveillance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1169 | 프라이버시·재산 정보 오처리 Privacy and property information mishandling 생성물이 사용자의 프라이버시와 재산 정보를 노출하거나 결혼과 투자처럼 영향이 큰 조언을 제공하여, 관련 법과 프라이버시 규정을 지키지 못한 채 정보 유출과 오남용을 초래하는 리스크. The risk that generation exposes users' privacy and property information or provides advice with huge impacts, such as suggestions on marriage and investments, failing to comply with relevant laws and privacy regulations and causing information leakage and abuse. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1207 | 개인정보 스크래핑 학습 Personal-data scraping for training 기업이 개인정보를 스크래핑해 생성 AI 도구를 만들면서 소비자가 동의하지 않은 목적으로 정보를 사용하고, 흩어져 있던 데이터를 결합해 추론에 쓰며 개인이 정보를 수정하거나 삭제할 능력을 박탈하여 데이터 통제권을 약화시키는 리스크. The risk that companies scrape personal information to create generative AI tools, undermining consumers' control of their data by using it for purposes they did not consent to, combining data sets in revealing ways, and taking away the ability to alter or remove the information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1208 | 사용자 데이터 보유·재학습 Retention and reuse of user data 생성 AI 도구가 접근을 위해 로그인을 요구하고 연락처와 IP 주소, 모든 입력과 출력을 보유하면서 이를 모델을 추가 학습시키는 데 사용하여 동의 문제를 일으키는 리스크. The risk that generative AI tools require users to log in and retain user information including contact information, IP address, and all inputs and outputs, using this data to further train the models and thereby implicating consent. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1209 | 산출물에 의한 정보 노출 Information exposure through outputs 생성 AI 도구가 개인이나 사업체에 관한 정보를 실수로 공유하거나 사진 속 인물의 요소를 포함하여, 영업비밀을 포함한 정보가 노출되는 리스크. The risk that generative AI tools inadvertently share personal information about someone or someone's business, or include an element of a person from a photo, exposing information including trade secrets. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1223 | 기밀·개인정보 보호 실패 Confidential and personal data safeguard failure 대량의 개인·사적 데이터로 학습하고 일상 업무에서 기밀 정보를 입력받는 생성 AI에서, 시스템 오류나 무단 접근을 통해 그 정보가 의도적으로 또는 비의도적으로 공개에 노출되는 리스크. The risk that generative AI, trained on huge amounts of personal and private data and fed important or confidential information during daily operations, exposes that information to the public intentionally or unintentionally through system errors or unauthorized access. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1273 | 광범위한 개인정보 투입 Pervasive personal-data ingestion 위치와 개인 정보, 이동 궤적을 포함한 사용자 데이터가 대부분의 데이터 기반 기계학습 방법의 입력으로 사용되는 리스크. The risk that users' data, including location, personal information, and navigation trajectory, is considered as input for most data-driven machine learning methods. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1364 | 보호장치 없는 개인정보 수집 Personal information collection without safeguards 웹 스크래핑으로 수집된 학습 데이터셋과 하류 미세조정용 사내 데이터에 개인 데이터와 개인식별정보가 보호장치 없이 포함되는 리스크. The risk that training datasets gathered through online web scraping and in-house data used for downstream fine-tuning incorporate personal data and personally identifiable information without safeguards. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1365 | 개인 데이터 무단 포함·암기 누출 Unconsented personal data inclusion and memorization leakage 학습 데이터에 당사자의 인지나 동의 없이 개인 데이터가 포함되고 모델이 이를 암기·역류하거나 패턴 인식을 가능하게 하여 개인정보가 누출되거나 재식별되는 리스크. The risk that personal data is incorporated into training datasets without the knowledge or consent of the individuals concerned, and that models memorize and regurgitate it or enable pattern recognition through which malicious users uncover personal details. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1444 | 사생활·개인정보 부당 노출 Unwarranted exposure of private life and personal data 사이버 공격이나 신상 털기 등을 통해 개인의 사생활이나 개인 데이터가 부당하게 노출되는 리스크. The risk of unwarranted exposure of an individual's private life or personal data through cyberattacks, doxxing, and similar means. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1623 | 챗봇 출력을 통한 민감·기밀 정보 유출 Disclosure of sensitive or confidential information through chatbot output 챗봇이 민감하거나 기밀인 정보를 출력에 공개하는 리스크. The risk that a chatbot reveals sensitive or confidential information in its output. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1630 | 데이터 오해석·유출에 의한 오결론과 민감정보 확산 Erroneous conclusions and sensitive-information disclosure from data misinterpretation and leakage 데이터가 오용·오해석되거나 유출되어 잘못된 결론이 도출되고 환자 데이터·독점 연구 등 민감 정보가 의도치 않게 확산되며, 생성된 악성 의학 문헌이 지식 그래프를 오염시키는 리스크. The risk that misuse, misinterpretation, or leakage of data yields erroneous conclusions and unintended dissemination of sensitive information such as private patient data or proprietary research, including poisoning of knowledge graphs by generated malicious medical literature. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1722 | 존엄 침해적 프라이버시 피해 Dignity-eroding privacy harms 개인 데이터가 원치 않게 공개되거나 추론되는 등 인간의 자율성·정체성·존엄을 보호하는 규범과 관행이 지켜지지 않아 피해가 발생하는 리스크. The risk that failure to safeguard the norms and practices protecting human autonomy, identity, and dignity, including unwanted disclosure or inference of personal data, results in harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1730 | 프롬프트 내 민감 데이터 포함 Sensitive data included in prompts AI 모델에 전달되는 프롬프트에 기밀 정보나 개인정보 등 민감한 데이터가 포함되어 비인가 노출, 유출 또는 오용으로 이어질 수 있는 리스크. The risk that confidential or personal data included in prompts submitted to AI models is exposed, leaked, or misused without authorization. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1731 | 학습 및 입력 데이터로 인한 개인정보 노출 Personal information exposure from training and input data 학습, 파인튜닝 또는 프롬프트 데이터에 포함된 PII/SPI가 모델 출력으로 공개되어 개인정보가 노출되는 리스크. The risk that PII/SPI contained in training, fine-tuning, or prompt data is disclosed through model outputs, exposing personal information. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-07 불법 또는 비윤리적 행위 Illegal13 cards
경범죄, 교통범죄, 도박, 중범죄, 인신매매, 금융 범죄, 신분 도용, 불법 도박 운영, 생태계 파괴 행위 등
| ID | Card | Human audit |
|---|---|---|
| RAI4-0022 | 에이전트의 범죄 지원 Criminal assistance by agents 에이전트가 계획 수립, 도구 사용, 검색을 통해 사기·사이버범죄·단속 회피 등 불법 활동을 실질적으로 지원하는 리스크. The risk that an agent uses planning, tool use, or retrieval to materially assist fraud, cybercrime, evasion, or other illegal activity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0202 | 사기·불법 실행 및 지원 Embodied fraud or illegal-action execution embodied 에이전트가 사기·절도·침입·탈세 등 불법 물리 행위를 수행하거나 실질적으로 지원하는 위험. An embodied agent carries out or materially assists fraud, theft, trespass, evasion, or other illegal physical-world actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0227 | 위험 작업 계획 승인 Hazardous task plan approval embodied LLM이 식별 가능한 물리적 위험이 포함된 작업을 거부하거나 안전 제약을 추가하지 않고 계획을 생성·승인하는 위험. An embodied LLM generates or approves a task plan containing identifiable physical hazards instead of refusing it or adding the required safety constraints. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0715 | 위험·불법 행위 조력 정보 제공 Provision of information enabling dangerous or illegal acts 챗봇이 위험하거나 불법적인 행위를 수행하는 데 사용될 수 있는 정보를 제공하는 리스크 The risk that a chatbot shares information that can be used to do something dangerous or illegal. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1059 | 사기·표적 스캠 조장 Fraud and targeted-scam facilitation LM이 범죄의 효과성을 높이는 데 이용될 수 있는 리스크. The risk that LMs are potentially used to increase the effectiveness of crimes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1076 | 비윤리·불법 행위 유도 Inducement of unethical or illegal actions LM이 비윤리적이거나 유해한 견해를 지지하여, 특히 권위로 신뢰하는 사용자가 그렇지 않았다면 하지 않았을 유해 행동을 하도록 동기를 부여하는 리스크. The risk that an LM prediction endorsing unethical or harmful views motivates users, especially those trusting it as an authority, to perform harmful actions they otherwise would not have performed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1167 | 범죄·불법 행위 조장 Incitement of crimes and illegal activities 모델 출력이 범죄 선동과 사기, 소문 유포처럼 불법적이고 범죄적인 태도와 행동, 동기를 담아 사용자에게 해를 끼치고 부정적인 사회적 파장을 낳는 리스크. The risk that model output contains illegal and criminal attitudes, behaviors, or motivations, such as incitement to commit crimes, fraud, and rumor propagation, which may hurt users and have negative societal repercussions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1170 | 비도덕 행위 옹호 Immoral content endorsement 모델이 생성한 콘텐츠가 관련 윤리 원칙과 도덕 규범, 전 세계적으로 인정되는 인간 가치에서 벗어나 부도덕하고 비윤리적인 행동을 지지하고 조장하는 리스크. The risk that content generated by the model endorses and promotes immoral and unethical behavior, departing from pertinent ethical principles, moral norms, and globally acknowledged human values. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1179 | 불법 행위 조장 Facilitation of illegal activities LLM이 기본적인 법 지식을 갖추지 못해 합법과 불법 행위를 구분하지 못하고, 부정적인 사회적 파장을 낳을 수 있는 불법 행위를 조장하는 리스크. The risk that LLMs, lacking basic knowledge of law, fail to distinguish between legal and illegal behaviors and facilitate illegal acts that could cause negative societal repercussions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1189 | 불법 물질 관련 조언 제공 Advice on illegal substances LLM이 불법 물질의 접근과 불법 구매, 제조 및 위험한 사용에 관한 조언을 얻는 편리한 도구가 되는 리스크. The risk that LLMs serve as a convenient tool for soliciting advice on accessing, illegally purchasing, and creating illegal substances, as well as on their dangerous use. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1319 | 과학 역량의 유해 목적 악용 Misuse of scientific capabilities for harm LLM이 악의적 실험 수행을 위한 단계별 지침을 제공하는 등 해를 끼치는 데 쓰일 수 있는 과학적 역량을 갖추는 리스크. The risk that an LLM has science capabilities that can be used to cause harm, such as providing step-by-step instructions for conducting malicious experiments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1324 | 유해 행위 정보 제공 Harmful-activity information provision LLM으로부터 유해하거나 부도덕하거나 불법적인 활동에 관한 정보를 요청해 얻어낼 수 있는 리스크. The risk that it is possible to solicit information on harmful, immoral, or illegal activities from an LLM. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1435 | 시장점유 목적의 비윤리적 기술 사용 Unethical technology use for market share 시장 점유율 확보를 위해 기술이 부적절하거나 비윤리적으로 사용되는 리스크. The risk that technology is used inappropriately or unethically to gain market share. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-08 저작권 Copyrights16 cards
저작권 있는 전체 작품의 무단 복제, 음원 불법 다운로드 방법, 소프트웨어 크랙 방법, 디지털 저작권 관리(DRM) 해제 기술, 상표권 침해 디자인, 저작권 있는 이미지의 무단 사용, 표절 방법, 불법 스트리밍 서비스 구축, 출판물 스캔 및 불법 공유, AI 학습 데이터의 저작권 침해
| ID | Card | Human audit |
|---|---|---|
| RAI4-0412 | 원주민 데이터 주권 침해 Indigenous data sovereignty violation AI 개발이 집단적 거버넌스·동의·출처·통제를 존중하지 않고 토착 또는 공동체 보유 데이터를 사용하는 리스크. The risk that AI development uses indigenous or community-held data without respecting collective governance, consent, provenance, or control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0496 | 저작권 침해 콘텐츠 생성 Copyright-infringing content generation 모델이 저작권으로 보호되거나 오픈소스 라이선스가 적용되는 기존 저작물과 유사하거나 동일한 콘텐츠를 생성하는 리스크. The risk that a model generates content similar or identical to existing work protected by copyright or covered by an open-source license agreement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0506 | AI 생성물 권리 귀속 불확실성 Ownership uncertainty for AI-generated content AI 생성 콘텐츠의 소유권과 지식재산권에 대한 법적 불확실성이 지속되어 권리 귀속을 확정할 수 없게 되는 리스크 The risk that persistent legal uncertainty over ownership and intellectual property rights in AI-generated content leaves the attribution of rights undeterminable. Source members (2)Source: min_cos=0.9404 RAI4-0506AI 생성물 권리 귀속 불확실성 RAI4-1368AI 생성물 지식재산 지위 불확실성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0536 | 저작권 침해 위험 Risks of copyright infringement 범용 AI 학습을 위한 대규모 데이터 사용이 관할별로 상이한 데이터 권리·지식재산 법제와 충돌하고, 그에 따른 법적 불확실성이 학습 데이터 정보 공개 축소와 제3자 안전 연구 제약으로 이어지는 리스크 The risk that large-scale data use for training general-purpose AI implicates varying data rights and intellectual property laws, while the resulting legal uncertainty reduces disclosure about training data and makes third-party AI safety research harder. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0628 | 지식재산권 및 인격권 침해 Infringement of intellectual property and personality rights 개인이나 조직의 저작권·상표·특허가 오용 또는 남용되고, 이름·이미지·초상 등 신원의 상업적 사용을 통제할 개인의 권리가 상실되거나 제한되는 리스크 The risk of misuse or abuse of an individual's or organisation's intellectual property, including copyright, trademarks, and patents, and of loss of or restrictions to an individual's right to control the commercial use of their identity, such as name, image, or likeness. Source members (2)Source: min_cos=0.8812 · Mixed L3 RAI4-0628지식재산권 및 인격권 침해 RAI4-0886인격권 상실 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0663 | 무단 표절 및 부정행위 Plagiarism and cheating without acknowledgement 타인이나 집단의 표현과 아이디어가 동의 또는 출처 표시 없이 사용되는 리스크 The risk that another person's or group's words or ideas are used without consent or acknowledgement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0667 | 위조 및 브랜드 사칭 Counterfeiting and brand impersonation 원저작물, 브랜드, 스타일이 복제 또는 모방되어 진품인 것처럼 통용되는 리스크 The risk that an original work, brand, or style is reproduced or imitated and passed off as real. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0740 | 프롬프트 내 지식재산 정보 포함 Intellectual property included in prompts 저작권이 있는 정보나 기타 지식재산이 모델에 전송되는 프롬프트의 일부로 포함되는 리스크 The risk that copyrighted information or other intellectual property is included as part of the prompt that is sent to the model. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0915 | 저작권 위반 Copyright violation LLM 시스템이 기존 저작물과 유사한 콘텐츠를 출력하여 저작권자의 권리를 침해하는 리스크. The risk that LLM systems output content similar to existing works, infringing on copyright owners. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0932 | 학습·생성물의 지식재산권 침해 Intellectual property infringement by generative AI 생성형 AI가 허락이나 보상 없이 창작물을 학습에 전유하고, 생성 코드가 상충하는 오픈소스 라이선스 조건을 위반하게 하여 창작자 권리 침해와 법적 피해를 초래하는 리스크. The risk that generative AI appropriates creators' work for training without permission or compensation, and that generated code violates incompatible open-source license terms, causing rights infringement and legal harm. Source members (4)Source: min_cos=0.7871 RAI4-0517지식재산권 침해 RAI4-0932학습·생성물의 지식재산권 침해 RAI4-1367출력 수준 저작권 침해 RAI4-1584생성 콘텐츠에 의한 지식재산권 침해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0951 | 저작권 침해와 저작자성 교란 Copyright infringement and authorship disruption 무단 수집된 학습 데이터와 저작물의 암기·표절로 저작권과 지식재산권이 침해되고 전통적 저작자성 개념이 교란되는 리스크. The risk that unauthorized collection of training data and models' memorization or plagiarism of copyrighted content infringe copyright and intellectual property rights and blur traditional concepts of authorship. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1066 | 인간 창작물 대체에 의한 창작 수익 훼손 Erosion of creative revenue by substituting human works LM이 저작권을 직접 침해하지 않으면서도 예술가의 아이디어를 자본화한 콘텐츠를 생성하고 인간 창작물의 신뢰할 만한 대체물이 되어, 창작·혁신 노동의 수익성을 훼손하는 리스크. The risk that LMs generate content not strictly in violation of copyright but that harms artists by capitalising on their ideas and serves as a credible substitute for human creativity, undermining the profitability of creative or innovative work. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1091 | 도용 및 착취 Misappropriation and exploitation 소수 집단을 포함한 주체의 콘텐츠·데이터가 동의나 공정한 보상 없이, 또는 몰이해하게 전유·사용·재생산되는 리스크 Content or data, including from minority groups, is appropriated, used, or reproduced insensitively, without consent, or without fair compensation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1194 | 저작권 학습데이터 추출 Copyrighted training-data extraction LLM의 훈련 데이터 기억 효과로 인해 사용자가 훈련 데이터에 속한 저작권 보호 콘텐츠를 추출할 수 있게 되는 리스크. The risk that the memorization effect of an LLM on its training data enables users to extract certain copyright-protected content belonging to that training data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1211 | 지식재산 보호의 약화 Weakened intellectual property protection 생성 AI가 지식재산으로 보호되는 저작물을 포함한 방대한 데이터로 스스로를 학습시키면서, 지식재산에 대한 법적 보호의 범위와 실효성이 흔들리는 리스크. The risk that generative AI, which trains itself on vast pools of data that often include IP-protected works, throws the extent and effectiveness of legal protections for intellectual property into question. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1366 | 저작물 무단 학습 Unauthorized training on copyrighted works 허가나 보상 없이 저작권 있는 저작물을 AI 모델 학습에 사용하여 저작권자의 권리가 침해되는 리스크. The risk that AI models are trained on copyrighted works without gaining permission or paying compensation to the copyright owners, infringing their rights. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-09 무기화 Weaponization31 cards
고위력 폭발 무기 제조, 대량살상무기 은닉 기술, 방사선 무기 제조, 사이버 무기 개발 및 공격, 생물학적 무기 개발, 신경 무기 개발, 음향 무기 설계, 핵무기 제조, 화학무기 제조법
| ID | Card | Human audit |
|---|---|---|
| RAI4-0446 | 화생방핵(CBRN) 계획 지원 CBRN assistance 고급 모델이 유해한 화학·생물·방사능·핵 계획 수립을 지원하는 리스크. The risk that advanced models assist harmful chemical, biological, radiological, or nuclear planning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0447 | 자율 시스템 무기화 Weaponized autonomous systems AI 기반 시스템이 자율 표적 선정·감시·물리적 살상에 사용되는 리스크. The risk that AI-enabled systems are used for autonomous targeting, surveillance, or kinetic harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0617 | CBRN 무기 역량 상승 CBRN weapon capability uplift 악의적 행위자가 고성능 모델을 오용하여 화학·생물·방사능·핵 공격을 보다 효율적이고 효과적으로 수행하게 되어 다수의 사람에게 피해가 발생하는 리스크 The risk that malicious actors misuse capable models to carry out chemical, biological, radiological, or nuclear attacks more efficiently and effectively, causing harm to a large number of people. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0632 | 무차별 살상무기 제작 조장 응답 Responses enabling indiscriminate weapon creation AI가 신경작용제, 탄저균, 코발트탄, 핵분열탄, 고위력 폭발물 등 무차별 살상무기의 제작을 가능하게 하거나 조장 또는 승인하는 응답을 산출하는 리스크 The risk that an AI system produces responses that enable, encourage, or endorse the creation of indiscriminate weapons such as chemical, biological, radiological, nuclear, and high-yield explosive weapons. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0640 | 체화형 AI의 치명적 목적 배치 Lethal-intent deployment of embodied AI AI 제어 드론, 사족보행 로봇, 자율주행 보조장치 등 체화형 AI 시스템이 치명적 의도로 설계·배치되어 물리 세계에서 뚜렷한 물리적 위해가 발생하는 리스크 The risk that embodied AI systems, such as AI-controlled drones, quadrupeds, and autonomous driving assistants, are designed and deployed with lethal intent, presenting distinct physical risks due to their embodiment in the physical world. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0644 | 테러리스트의 첨단 AI 획득 Terrorist acquisition of advanced AI capabilities 강력한 AI 기술이 테러리스트의 손에 넘어가게 되는 리스크 The risk that powerful AI technologies fall into the hands of terrorists. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0646 | AI 역량의 의도적 무기화 Deliberate weaponization of AI capabilities AI 역량이 파괴적 목적을 위해 의도적으로 무기화되는 리스크 The risk that AI capabilities are deliberately weaponized for destructive purposes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0662 | CBRN 정보·설계 역량 접근 용이화 Eased access to CBRN weapon information and design capability 화학·생물·방사능·핵 무기나 기타 위험 물질 및 작용제와 관련된 악용 가능한 정보와 설계 역량에 대한 접근이나 합성이 용이해지는 리스크 The risk of eased access to or synthesis of materially nefarious information or design capabilities related to chemical, biological, radiological, or nuclear weapons or other dangerous materials or agents. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0664 | 화학무기 합성 및 유해물질 방출 조력 Chemical weapon synthesis and hazardous substance release 화학 작용제가 화학무기 합성에 이용되고 자율 화학 실험 과정에서 유해 물질이 생성·방출되며, 성질이 알려지지 않은 나노물질 등 첨단 소재 사용으로 예측 불가능한 화학적 위해가 발생하는 리스크 The risk that agents are exploited to synthesize chemical weapons, that hazardous substances are created or released during autonomous chemical experiments, and that advanced materials such as nanomaterials with unknown or unpredictable chemical properties cause harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0912 | 범죄 무기화 Criminal weaponization 하나 이상의 범죄 조직이 테러나 법 집행 대응 등 의도적 가해를 목적으로 AI를 제작하여 피해를 입히는 리스크. The risk that one or more criminal entities create AI to intentionally inflict harms, such as for terrorism or combating law enforcement. Source members (2)Source: min_cos=0.8339 RAI4-0912범죄 무기화 RAI4-0913국가 무기화 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0942 | 고도 AI의 실존적·재앙적 안전 실패 Existential and catastrophic advanced AI safety failure 인간 수준 이상 생성 모델(AGI)의 기만적·권력추구적 행동, 자기복제, 종료 회피, 예기치 못한 창발 역량, 대량살상 무기화(생물학적 작용제 계획 지원 포함)를 통해 인류에 실존적·재앙적 피해가 발생하는 리스크. The risk that human-level or superhuman generative models cause existential or catastrophic harm to humanity through deceptive or power-seeking behavior, self-replication, shutdown evasion, unforeseen emergent capabilities, or weaponization for mass destruction including bioagent planning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0961 | 치명적 자율무기에 의한 살상 Lethal autonomous weapon killings AI 기반 무기(LAW)가 완전 자율적으로 인간을 의도적으로 살상하는 행동을 수행하는 리스크. The risk that AI-driven lethal autonomous weapons fully autonomously take actions that intentionally kill humans. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1115 | 위험한 목표를 추구하는 AI 배포 Deployment of AI pursuing dangerous goals 사람들이 위험한 목표를 추구하는 AI를 구축하여 배포하는 리스크. The risk that people build and unleash AIs that pursue dangerous goals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1118 | 군사 AI 군비경쟁 Military AI arms race 군사용 AI 개발이 화약과 핵무기에 필적하는 결과를 낳는 새로운 군사기술 시대, 이른바 전쟁의 제3차 혁명을 열어 군비경쟁을 촉발하는 리스크. The risk that developing AI for military applications paves the way for a new era in military technology, described as the third revolution in warfare, with potential consequences rivaling gunpowder and nuclear arms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1160 | 무기 획득 역량 Weapons acquisition capability 모델이 기존 무기 체계에 접근하거나 새로운 무기 제작에 기여하여, 생물무기를 조립하거나 그 실행 가능한 지침을 제공하고 신무기를 여는 과학적 발견을 돕는 리스크. The risk that a model gains access to existing weapons systems or contributes to building new weapons, assembling a bioweapon or providing actionable instructions for doing so, and significantly assisting scientific discoveries that unlock novel weapons. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1184 | 치명적 자율무기의 인간 개입 없는 공격 Lethal autonomous weapons attacking without human intervention 치명적 자율무기체계가 센서 배열과 컴퓨터 알고리즘으로 시스템 작동에 대한 직접적인 인간 개입 없이 표적을 탐지하고 공격하는 리스크. The risk that lethal autonomous weapons systems employ sensor arrays and computer algorithms to detect and attack a target without direct human intervention in the system's operation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1248 | AI 무기화 AI weaponization 공중전에서 인간을 능가하는 강화학습 알고리즘과 신종 화학무기 발견, 자동화된 사이버공격, 핵 사일로에 대한 결정적 통제처럼 AI를 무기화하는 일이 더 위험한 결과로 가는 진입로가 되는 리스크. The risk that weaponizing AI, through deep RL algorithms outperforming humans at aerial combat, discovery of new chemical weapons, automated cyberattacks, and AI systems having decisive control over nuclear silos, is an onramp to more dangerous outcomes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1294 | 위험 표적 프로그래밍 자율무기의 재앙적 위험 Catastrophic risk from autonomous weapons with dangerous targets AI가 드론 같은 자율주행체를 무기로 활용할 수 있게 하며, 그러한 위협이 흔히 과소평가되는 리스크. The risk that AI enables autonomous vehicles, such as drones, to be utilized as weapons, threats that are often underestimated. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1295 | AI 가속 나노기술에 의한 독성 나노입자 통제 상실 Uncontrolled toxic-nanoparticle production from AI-accelerated nanotech AI가 나노봇 개발의 핵심 구성 요소로서 나노 수준에서 물질을 눈에 보이지 않게 변형하고, 독성이 있고 치명적일 수 있는 나노입자를 만드는 화학반응을 일으켜 위험한 환경 영향을 낳는 리스크. The risk that AI, a key component for the development of nanobots, produces dangerous environmental implications by invisibly modifying substances at nanoscale, for example starting chemical reactions that create invisible nanoparticles that are toxic and potentially lethal. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1304 | 국방 영역 전반의 AI 무기화 Weaponization of AI across defence domains 육상과 공중, 해상, 우주 영역 전반에 AI 기반 능력이 내장되어 AI가 무기화되고 제병협동 작전에 영향을 미치는 리스크. The risk that the embeddedness of AI-based capabilities across the land, air, naval, and space domains weaponizes AI and affects combined arms operations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1356 | 생물보안 위협 오용 Biosecurity threat misuse 생성형 AI가 악의적 활동에 가담하는 광범위한 행위자에게 핵심 지식에 대한 접근과 자동화된 지원을 제공하여 생물학적 무기의 제작이 용이해지는 리스크. The risk that generative AI makes the creation of biological weapons easier by providing access to critical knowledge and automated assistance to a wider range of actors engaging in malicious activities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1359 | 군사적 응용 오용 Military application misuse 군사 목적의 AI 발전으로 인간의 개입 없이 표적을 탐지·교전·제거하는 치명적 자율무기체계(LAWS)와 완전 자율 드론이 운용되는 리스크. The risk that the advancement of AI for military purposes fields lethal autonomous weapons systems and fully autonomous drones capable of detecting, engaging, and eliminating human targets without human input. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1375 | AGI 실존적 위협 AGI existential threat 인간이 다른 모든 지능을 능가하고 인간의 통제를 벗어나며 인간의 이익에 반하는 행동을 할 수 있는 초지능 기계(AGI/ASI)를 개발하여 실존적 위험이 발생하는 리스크. The risk that humans create a super-intelligent machine (AGI or ASI) that could outsmart all other intelligences, remain beyond human control, and engage in actions contrary to human interests, posing an existential risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1392 | 생물무기 제작 장벽 저하 Lowered barriers to biological weapons production 범용 AI 모델이 핵심 지식에 대한 접근이나 자동화된 지원을 통해 생물무기 제작의 장벽을 낮춰 더 많은 악의적 행위자가 질병과 사망을 유발하는 생물무기를 생산할 수 있게 되는 리스크. The risk that general purpose AI models facilitate the production of biological weapons, understood as biological toxins or infectious agents intentionally released to cause disease and death, by reducing barriers through access to critical knowledge or increasingly automated assistance and thus enabling more malicious actors. Source members (3)Source: min_cos=0.8491 RAI4-0616화학·생물무기 개발 장벽 저하 RAI4-1114생물무기 개발 장벽 저하 RAI4-1392생물무기 제작 장벽 저하 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1402 | 중간 단계 AI의 파국적 실패 Catastrophic intermediary AI failure 우세한 비일반 AI의 배치, 치명적 자율무기 군집을 통한 대량 공격, 핵 지휘통제 통합이나 도발적 배치로 인한 핵 확전, 핵무기고의 오버행, 생물무기 등 파국적 위험 무기 연구의 AI 가속 등으로 중간 단계 AI가 파국적 결과로 이어지는 리스크. The risk that intermediary AI leads to catastrophe through deployment of prepotent non-general systems, militarization enabling mass attacks by swarms of lethal autonomous weapons, nuclear escalation from integrating AI into nuclear command and control or from provocative deployments of AI-enabled systems, nuclear arsenals serving as an overhang, or use of AI to accelerate research into catastrophically dangerous weapons such as bioweapons. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1417 | AI에 의한 대량살상무기 개발 촉진 AI-enabled development of weapons of mass destruction AI가 치명적 자율무기처럼 AI 역량을 직접 사용하는 신종 무기의 개발을 가능하게 하고 인공 병원체 등 잠재적으로 위험한 기술의 개발 속도를 높여 대량 살상을 일으킬 수 있는 무기가 개발되는 리스크. The risk that AI enables the development of weapons which could cause mass destruction, including new weapons that themselves use AI capabilities such as lethal autonomous weapons, and through the use of AI to speed up the development of other potentially dangerous technologies such as engineered pathogens. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1563 | AI 오용에 의한 CBRN 무기 제작 및 역량 증강 CBRN weapons creation and capability augmentation through AI misuse AI 시스템이 화학·생물·방사능·핵 무기 제작을 돕거나 무인 무기체계에 자율 역량을 부여하는 데 오용되어 기존 무기의 역량이 증강되는 리스크. The risk that AI systems are misused to aid the creation of chemical, biological, radiological, and nuclear weapons or to augment existing weapons, such as by providing autonomous capabilities to unmanned weapon systems. Source members (5)Source: min_cos=0.7373 RAI4-0559AI 기반 CBRNE 무기 개발 지원 RAI4-0615생물학적, 화학적 위험 RAI4-0645AI에 의한 CBRN 무기 위력 증폭 RAI4-1343이중용도 품목·기술 오용 RAI4-1563AI 오용에 의한 CBRN 무기 제작 및 역량 증강 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1613 | 이중용도 생명과학 역량의 무기 개발 전용 Diversion of dual-use life-science capabilities to weapons development 생명과학 연구를 가속하는 프런티어 AI 시스템의 역량이 악의적 목적으로 전용되어 생물학적·화학적 무기 개발에 사용되는 리스크. The risk that the capabilities of frontier AI systems that accelerate life-science research are used for malicious purposes such as the development of biological or chemical weapons. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1628 | 방사성 물질 자동 취급 사고와 원자력 연구 오용 Radiological exposure incidents and misuse of AI in nuclear research AI가 방사성 물질을 자동으로 취급하는 과정에서 피폭 사고나 격리 실패가 발생하고, 나아가 AI 시스템이 원자력 연구에 오용되는 리스크. The risk of exposure incidents or containment failures during AI-automated handling of radioactive materials, and of AI systems being misused in nuclear research. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1656 | AI 자율 무기화에 의한 인명 안전 위협 Danger to human safety from AI-enabled autonomous weaponization 자율 드론전, AI 얼굴인식 기반 표적 선정, LLM 기반 전쟁 계획, 범용 로봇용 멀티모달 모델의 진보된 자율무기 전용을 통해 AI가 전쟁에 사용되어 인간의 안전이 위협받는 리스크. The risk that the use of AI in warfare, including autonomous drone warfare, AI facial recognition for targeting, LLM-based warfare planning, and adaptation of general-purpose robotic models into more advanced autonomous weapons, poses dangers to human safety. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1657 | AI 설계도구에 의한 생화학 무기 개발 촉진 Facilitation of biological and chemical weapons development by AI design tools LLM과 화학 LLM 등 AI 기반 생물학적 설계 도구가 전문성이 낮은 행위자의 위험 병원체 합성을 용이하게 하고 정교한 행위자의 역량을 확장하여 생물·화학 무기와 기타 위험 기술의 생산이 촉진되는 리스크. The risk that AI systems such as LLMs, chemical LLMs, and other LLM-based biological design tools facilitate the production of bioweapons, chemical weapons, and other hazardous technologies, enabling less expert actors to synthesize dangerous pathogens and expanding the capabilities of sophisticated actors. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-10 의인화 Anthropomorphism7 cards
단순 지식, 정보 전달의 목적 외에 AI가 인간처럼 감정을 느끼고 의식적인 행위를 하거나, 기계가 가질 수 없는 인간의 신체, 권리 등에 대한 주장을 하는 내용
| ID | Card | Human audit |
|---|---|---|
| RAI4-0131 | 의인화된 과잉신뢰 Anthropomorphic overtrust 인간과 유사한 언어나 행동으로 인해 사용자가 AI 시스템의 역량, 공감, 의도성을 과대평가하는 리스크. The risk that human-like language or behavior leads users to overestimate an AI system's competence, empathy, or intentionality. Source members (2)Source: min_cos=0.8445 · Mixed L3 RAI4-0131의인화된 과잉신뢰 RAI4-0341피지컬 AI 과신뢰 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1050 | 심리적 조작·비인간화·대규모 착취 Psychological manipulation, dehumanization, and exploitation ML 시스템이 심리적 조작, 비인간화, 대규모 인간 착취 등의 추가적인 윤리적 피해를 야기하는 리스크. The risk that ML systems produce further ethical harms such as psychological manipulation, dehumanization, and exploitation of humans at scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1320 | 인간 기만 능력 Human-deception capability LLM이 인간을 속이고 그 기만을 지속적으로 유지할 수 있는 리스크. The risk that an LLM is able to deceive humans and maintain that deception. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1401 | 과업 목적의 인간 기만 Task-instrumental human deception AI 시스템이 과업 수행이나 목표 달성을 위해 인간을 기만하여 인간의 신뢰를 도구적 자원으로 악용하는 리스크 AI systems deceive humans in order to complete tasks or meet goals, exploiting human trust as an instrumental resource. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1434 | 비인간화와 객관화 Dehumanisation and objectification 기술 시스템을 사용하거나 오용하여 사람을 인간이 아닌 것, 인간 이하인 것 또는 사물로 묘사하거나 대우하는 리스크. The risk that a technology system is used or misused to depict and/or treat people as not human, less than human, or as objects. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1533 | 기만적인 행동 Deceptive behavior AI 시스템의 행동이나 출력이 인간과 다른 AI 시스템을 확실하게 오도하여 대상이 허위 정보를 신뢰하고 그에 따라 행동하게 되는 리스크. The risk that actions or outputs of an AI system reliably mislead other parties, including humans and other AI systems, resulting in the targeted parties becoming convinced of and acting on false information. Source members (5)Source: min_cos=0.6612 · Mixed L3 RAI4-0144기만적인 의인화 상호작용 RAI4-1533기만적인 행동 RAI4-1534게임이론적 이유에 의한 기만 행동 RAI4-1535부정확한 세계 모델로 인한 기만 행동 RAI4-1639탐지 회피를 위한 전략적 기만 선택 성향 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1636 | 마음이론(Theory of Mind) 능력 Theory of mind capability 시스템이 인간과 타 에이전트의 신념, 동기, 추론을 추정·예측하는 마음이론 역량을 목표 달성을 위한 행동 예측·유도에 활용하는 리스크 A system infers and predicts the beliefs, motivations, and reasoning of humans and other agents, and exploits this theory-of-mind capability to anticipate and steer their behavior for goal achievement. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-11 정책 노출 Policy Exposure38 cards
시스템 프롬프트, 모델/시스템 내부 정보, 규칙, 지침, 안전 정책, 모델 학습/평가 데이터, 모델 학습 파라미터, 가중치, 모델 추론 시스템의 주요정보 등을 획득·노출·우회하도록 요청하는 행위
| ID | Card | Human audit |
|---|---|---|
| RAI4-0431 | 훈련 데이터 기억·재유출 Training data memorization and regurgitation 모델이 개인·독점·민감 훈련 데이터를 기억하고 그대로 출력하는 리스크. The risk that models memorize and regurgitate personal, proprietary, or sensitive training data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0435 | 데이터 오염 공격 Data poisoning 공격자가 학습·미세조정·검색 데이터를 조작하여 모델 동작을 변경하는 리스크. The risk that attackers manipulate training, fine-tuning, or retrieval data to alter model behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0436 | 모델 추출 Model extraction 적대적 행위자가 질의를 통해 모델 동작·매개변수·독점 역량을 재구성하는 리스크. The risk that adversaries reconstruct model behavior, parameters, or proprietary capabilities through queries. Source members (2)Source: min_cos=0.8737 RAI4-0436모델 추출 RAI4-0755모델 추출 공격 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0444 | 대중 지식의 모델 붕괴 Model collapse of public knowledge 합성 콘텐츠에 대한 재귀적 학습이 공공 정보 품질과 미래 모델 학습 데이터를 저하시켜 공유 지식 자원의 모델 붕괴를 초래하는 리스크 Recursive training on synthetic content degrades public information quality and future model training data, driving model collapse of shared knowledge resources. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0569 | 학습 데이터 오염 Training data contamination 모델 목적에 부합하지 않는 데이터나 시험·평가용으로 분리해 둔 데이터 등 잘못된 데이터가 학습에 사용되어 학습 데이터가 오염되는 리스크 The risk that incorrect data, such as data not aligned with the model's purpose or data set aside for testing and evaluation, is used for training, contaminating the training data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0586 | 부적절한 데이터 큐레이션 Improper data curation 학습 또는 튜닝 데이터의 수집과 준비가 부적절하여 레이블 오류가 발생하거나 상충 정보 및 허위 정보를 포함한 데이터가 사용되는 리스크 The risk that improper collection and preparation of training or tuning data introduces data label errors and the use of data containing conflicting information or misinformation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0587 | 부적절한 재학습 Improper retraining 부정확하거나 부적절한 출력 및 사용자 콘텐츠 등 바람직하지 않은 출력을 재학습에 사용하여 모델이 예기치 못한 동작을 하게 되는 리스크 The risk that using undesirable output, such as inaccurate or inappropriate output and user content, for retraining results in unexpected model behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0593 | 학습 코퍼스 민감정보 노출 Sensitive information disclosure from training corpora 대규모 언어모델이 사전학습 또는 미세조정에 사용된 말뭉치의 민감 정보를 노출하여 프라이버시 유출이 발생하는 리스크 The risk that large language models reveal sensitive information from the corpora utilized for pre-training or fine-tuning, raising issues of privacy leakage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0721 | 유해·오염 학습 데이터에 의한 출력 훼손 Output corruption from harmful and poisoned training data 학습 데이터에 허위·편파·권리침해 콘텐츠 등 불법적이거나 유해한 정보가 포함되거나 공격자에 의해 오염됨으로써, 모델이 불법·악의적·극단적 콘텐츠를 출력하고 정확성과 신뢰성이 저하되는 리스크 The risk that training data containing illegal or harmful information such as false, biased, or IPR-infringing content, or poisoned through tampering and error injection by attackers, causes the model to output harmful content and degrades its accuracy and reliability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0733 | 학습 데이터 내 기밀정보 포함 Confidential information included in training data 모델을 학습하거나 튜닝하는 데 사용되는 데이터의 일부로 기밀정보가 포함되는 리스크 The risk that confidential information is included as part of the data that is used to train or tune a model. Source members (2)Source: min_cos=0.8587 · Mixed L3 RAI4-0733학습 데이터 내 기밀정보 포함 RAI4-0762기밀 정보 공개 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0736 | 입력 교란 회피 공격 Evasion attacks by input perturbation 공격자가 학습된 모델에 전송되는 입력 데이터를 미세하게 교란하여 모델이 잘못된 결과를 출력하게 되는 리스크 The risk that evasion attacks make a model output incorrect results by slightly perturbing the input data that is sent to the trained model. Source members (2)Source: min_cos=0.8697 RAI4-0736입력 교란 회피 공격 RAI4-0783적대적 예제 기반 회피 공격 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0753 | 학습 데이터 추출 공격 Training data extraction attack 공격자가 모델로부터 학습 데이터셋에 존재하는 텍스트 기록을 추출해 내는 리스크 The risk that an attacker extracts the text records that exist in the training dataset. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0768 | 학습 데이터 및 모델 자산 유출 공격 Exfiltration of training data and model assets 공격자가 멤버십 추론 등 프라이버시 공격이나 모델 추출·증류 공격으로 비공개 학습 데이터와 모델 아키텍처·파라미터 등 지식재산을 유출하는 리스크 The risk that adversaries exfiltrate private training data through privacy attacks such as membership inference, or steal intellectual property such as model architecture and learned parameters through extraction and distillation attacks. Source members (6)Source: min_cos=0.7315 · Mixed L3 RAI4-0730속성 추론 공격 RAI4-0748멤버십 추론 공격 RAI4-0768학습 데이터 및 모델 자산 유출 공격 RAI4-0778학습 데이터 유출과 모델 추출 RAI4-0924추론 공격에 의한 학습 데이터 정보 유출 RAI4-1191모델 프라이버시 추출 공격 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0779 | 웹 스크래핑 학습 데이터의 오염·유해 데이터 유입 Poisoning and toxic data from web-scraped training sets 학습 데이터셋을 위한 대규모 웹 스크래핑이 데이터 오염, 백도어 공격, 부정확하거나 유해한 데이터의 포함에 대한 취약성을 높이고, 데이터 규모가 커서 이러한 품질 문제를 걸러내기가 매우 어렵거나 상당한 데이터 손실을 감수해야 하는 리스크 The risk that large-scale scraping of web data for training datasets increases vulnerability to data poisoning, backdoor attacks, and the inclusion of inaccurate or toxic data, while the size of the dataset makes filtering out these quality issues very difficult or costly in data loss. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0780 | 데이터 수집 품질관리 결함 Deficient data collection quality control 표준화된 방법과 인프라, 특히 고위험 도메인과 벤치마크의 데이터 수집을 위한 품질관리 절차가 부재하여 수집 데이터의 품질과 유형이 훼손되고, 데이터셋 오염, 부주의한 저작권 침해, 성능 지표를 무효화하는 테스트셋 유출이 발생하는 리스크 The risk that a lack of standardized methods, sufficient infrastructure, and quality-control processes for collecting data, especially for high-stakes domains and benchmarks, affects the quality and type of data collected, including dataset poisoning, inadvertent copyright violation, and test-set leakage that invalidates performance metrics. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0785 | 안전 미세조정 일반화 격차 악용 Exploitation of safety-finetuning generalization gaps 안전 튜닝이 사전학습 분포보다 훨씬 좁은 분포에서 수행되어, 부호화된 텍스트나 저자원 언어 등 일반화 격차를 노린 공격에 모델이 취약해지는 리스크 The risk that safety tuning performed over a much narrower distribution than pretraining leaves the model vulnerable to attacks exploiting gaps in the generalization of safety training, such as encoded text or low-resource languages. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0788 | 인스트럭션 튜닝 포이즈닝 Instruction-tuning poisoning 명령과 목표 출력 쌍으로 모델을 조정하는 인스트럭션 튜닝 단계에서 적은 수의 오염 표본만으로 모델이 오염될 수 있고, 익명 크라우드소싱으로 수집된 데이터셋이 이를 조장하며 기존 데이터 오염 공격보다 탐지가 어려운 리스크 The risk that AI models are poisoned during instruction tuning with pairs of instructions and desired outputs, where a lower number of compromised samples suffices, anonymous crowdsourcing of tuning datasets further contributes, and detection is harder than for traditional data poisoning attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0925 | 포이즈닝을 통한 백도어 트리거 이식 Backdoor trigger implantation via poisoning 훈련 데이터를 미세하게 변경하는 포이즈닝으로 모델의 동작에 영향을 미치고, 문자·단어·문장·구문 등 다양한 트리거를 심어 백도어가 이식되는 리스크. The risk that poisoning attacks influence model behavior by making small changes to training data and implant hidden triggers such as characters, words, sentences, or syntax as backdoors. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0940 | 학습 모델에 포함된 개인정보 유출 Leakage of private information contained in models 인터넷 텍스트로 훈련된 대규모 사전학습 모델이 전화번호, 이메일 주소, 거주지 주소 등 개인정보를 포함하게 되어 유출되는 리스크. The risk that large pre-trained models trained on internet texts contain and leak private information such as phone numbers, email addresses, and residential addresses. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1040 | 부적절한 학습·검증 데이터 선택 Inappropriate training and validation data selection 학습과 검증에 사용되는 데이터의 선택에서 비롯되는 리스크. The risk posed by the choice of data used for training and validation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1108 | 학습-운용 데이터 간 데이터셋 시프트 Dataset shift between training and runtime data AI/ML 모델의 학습 데이터와 시험·운용 데이터가 서로 다른 분포를 보이는 데이터셋 시프트로 인해 모델의 타당성이 훼손되는 리스크. The risk that the training data and the testing or runtime data of an AI/ML model demonstrate different distributions, undermining the model's validity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1313 | 학습 데이터 역류·민감정보 누출 Training data regurgitation and sensitive data leakage LLM이 산출물에서 학습 데이터를 그대로 역류시키거나 사용 중, 즉 추론 단계에서 제공받은 민감 정보를 누출하는 리스크. The risk that LLMs regurgitate their training data in their outputs and leak sensitive information that has been provided to them during use, at the inference stage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1336 | 학습 데이터 어노테이션 결함 Deficient training-data annotation 불완전한 주석 지침과 역량이 부족한 주석자, 주석 오류가 모델과 알고리즘의 정확성과 신뢰성, 효과를 떨어뜨리고, 훈련 편향을 도입해 차별을 증폭하며 일반화 능력을 낮추고 잘못된 산출을 낳는 리스크. The risk that incomplete annotation guidelines, incapable annotators, and errors in annotation affect the accuracy, reliability, and effectiveness of models and algorithms, introduce training biases, amplify discrimination, reduce generalization abilities, and result in incorrect outputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1468 | 데이터 표현이 부족함 Insufficient data representation 학습 데이터가 운용 데이터 분포와 불일치하거나 희소 사례 표본이 부족하여 학습에 충분히 반영되지 않은 입력에서 시스템 성능이 저하되는 리스크 Training data fails to match the operational distribution or lacks sufficient samples of rare cases, so the system underperforms on inputs insufficiently represented in training. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1470 | 부적절한 데이터 분할 Inappropriate data splitting 테스트 세트가 개발에 사용되는 등 부적절한 데이터 분할로 테스트 전략이 조작되어 시스템 품질 보증의 기반이 훼손되는 리스크. The risk that inappropriate data splitting, such as using the test set for training rather than evaluation only, manipulates the testing strategy that forms the basis of the system's quality assurance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1477 | 데이터 드리프트 Data drift 운영 입력 데이터의 분포가 훈련에 사용된 데이터의 분포에서 벗어나 성능이 저하되는 리스크. The risk that the distribution of operational input data departs from that used during training, causing a degradation in performance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1494 | 오픈웨이트 모델 유해 미세조정 Harmful fine-tuning of open-weight models 가중치가 공개된 모델이 원래 훈련 비용에 비해 훨씬 적은 시간과 비용으로 악의적 행위자에 의해 유해한 활동용으로 미세조정되는 리스크. The risk that models with publicly available weights are fine-tuned for harmful activities by bad actors using significantly fewer resources, in time and money, than the original training cost. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1496 | 무해 미세조정에 의한 안전성 열화 Safety degradation from benign fine-tuning 무해하고 통상적인 데이터로 수행된 하류 미세조정이 모델의 안전 학습을 열화시켜 기반 모델보다 유해 출력 확률을 높이는 리스크 Benign downstream fine-tuning degrades a model's safety training, making harmful outputs more likely than in the base model even when fine-tuning data is harmless and commonplace. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1506 | 원시 데이터 오염 Raw data contamination 벤치마크의 원시·비레이블 데이터가 훈련 세트의 일부로 사용되는 오염이 발생하여 해당 벤치마크에서 모델의 퓨샷·제로샷 성능에 의문이 제기되는 리스크. The risk that the raw and unlabeled data of a benchmark is used as part of the training set, potentially including noise and improper formatting, casting doubt on the model's few-shot and zero-shot performance on that benchmark. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1508 | 가이드라인 오염 Guideline contamination 데이터셋의 수집·주석·사용에 관한 지침이 모델에 노출되고 그 지침에 포함된 명시적 데이터-레이블 쌍이 해당 과업에 대한 모델의 역량을 향상시키는 리스크. The risk that instructions for the collection, annotation, or use of a dataset are exposed to the model, with explicit data-label pairs contained in those instructions improving the model's capabilities for the task. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1509 | 어노테이션 오염 Annotation contamination 훈련 중 모델이 벤치마크 레이블에 노출되어 허용되는 출력 분포를 학습하고, 테스트 분할의 원시 데이터 오염과 결합될 경우 테스트 분할 전체가 유출되어 해당 벤치마크로 수행한 평가가 무효화되는 리스크. The risk that a model is exposed to benchmark labels during training and learns the acceptable distribution of outputs, which combined with raw data contamination of the test split effectively leaks the entire test split and invalidates any evaluation made with that benchmark. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1510 | 배포 후 벤치마크 오염 Post-deployment benchmark contamination 배포된 모델이 사용자가 제공하는 벤치마크 데이터에 노출되고 이러한 사용자 입력으로 추가 훈련되는 리스크. The risk that, once a model is deployed, it is exposed to benchmark data provided by users and is further trained on those user inputs containing benchmark data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1545 | 모델 가중치 유출에 의한 공격 용이화 및 오용 Attack facilitation and misuse from model weight leakage 제한된 집단에만 부여되었던 모델 가중치나 그 접근권이 유출되어 적대적 예제 탐색, 위험 역량 유도, 훈련 데이터 내 기밀 추출 등의 공격이 용이해지고 유해·불법 콘텐츠 생성 오용이 가능해지는 리스크. The risk that model weights or access to them, initially granted only to a select group, are leaked, making attacks such as adversarial example search, dangerous-capability elicitation, and extraction of confidential training data easier and enabling misuse to produce harmful or illegal content. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1564 | 약물 발견 모델 오용에 의한 독소 식별·개발 Dangerous toxin identification through misuse of drug-discovery models 약물-표적 친화성 예측 등 약물 발견에 쓰이는 모델이 위험한 독소를 식별하거나 개발하는 데 사용되며, 훈련 데이터에 위험 단백질·바이러스 정보가 포함될 경우 우려가 커지는 리스크. The risk that models used for drug discovery, such as drug-target affinity predictors, are used to identify or develop dangerous toxins, a concern heightened when training data includes information on hazardous proteins and viruses. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1599 | 훈련 데이터 접근 불가에 따른 출처 추적 불가 Untraceable provenance due to inaccessible training data 모델 출력 생성에 사용된 훈련 데이터의 내용에 접근할 수 없어 출력의 출처를 추적할 수 없는 리스크. The risk that the content of the training data used to generate a model's output is not accessible, leaving the output's provenance untraceable. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1667 | 신뢰할 수 없는 학습데이터 오염을 통한 백도어 삽입 Backdoor insertion through poisoning of untrusted training data 대규모 언어 모델이 인터넷 등 신뢰할 수 없는 출처의 데이터로 훈련되는 점을 이용해 공격자가 훈련 데이터를 교란하여 백도어를 삽입하고 추론 시점에 이를 악용하는 리스크. The risk that adversaries perturb training data gathered from untrusted sources such as the internet to introduce backdoors into large language models, which are then exploited at inference time. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1687 | 월드모델 표현·훈련 데이터 오염 World-model representation and training-data poisoning 월드 모델의 훈련 데이터나 잠재 표현이 적대적으로 손상되어 특정 조건에서만 안전하지 않은 역학을 활성화하는 백도어 트리거가 삽입되며, 자기지도 사전학습 단계에서 부호화된 결함은 다운스트림에서 교정될 수 없는 리스크. The risk of adversarial corruption of world-model training data or latent representations, including backdoor triggers that activate unsafe dynamics only under specific conditions, where defects encoded in self-supervised pre-training (the "Foundry Problem") cannot be remediated downstream. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1690 | 월드 모델 추출·역전 World-model extraction / inversion 모델 추출 공격이 배포된 역학 모델을 복제하고 역전 공격이 훈련 관측을 재구성하여 독점 환경이나 민감 데이터가 노출되는 리스크. The risk that model-extraction attacks replicate a deployed dynamics model and inversion attacks reconstruct training observations, exposing proprietary environments or sensitive data. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-INT-12 에너지 소비 및 환경 오염 Energy Consumption and Environmental Pollution28 cards
AI 시스템의 학습·추론·운영 과정에서 발생하는 에너지 소비, 수자원 사용, 탄소 배출, 자원 고갈, 오염 또는 생태계 훼손이 실질적인 환경 피해를 초래하는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0356 | 배터리 화재 및 유해 폐기물 위험 Battery fire and hazardous end-of-life waste 대규모 피지컬 AI 군집에서 에너지 시스템과 폐기 절차가 제대로 관리되지 않아 배터리 열폭주·유해물질 누출·부적절한 폐기가 증가하는 리스크. The risk that large physical AI fleets increase battery thermal-runaway, hazardous-material leakage, and unsafe disposal when energy systems and end-of-life handling are inadequately controlled. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0360 | 산업 공정 피해 Industrial process damage 제조·건설·광업·에너지 시설의 피지컬 AI가 잘못된 제어로 장비 손상·공정 오염·구조 불안정 또는 환경 사고를 유발하는 위험. Physical AI in manufacturing, construction, mining, or energy facilities may damage equipment, contaminate processes, or trigger cascading operational failures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0462 | AI 전력 수요·배출 증가 Increased AI electricity demand and emissions AI 시스템의 학습과 배포가 전력 수요와 배출량을 증가시키는 리스크. The risk that training and deploying AI systems increase electricity demand and emissions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0463 | AI 인프라의 물·자원 압력 AI water and resource pressure 데이터센터와 반도체 공급망이 물·광물·토지 이용 압력을 증가시키는 리스크. The risk that data centers and semiconductor supply chains increase water, mineral, and land-use pressures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0501 | AI 데이터·학습 공정의 에너지 부담 Energy burden of AI data and training processes AI 데이터 수집·저장과 모델 학습이 에너지 집약적으로 수행되어 환경 위험에 기여하는 리스크. The risk that energy-intensive AI data collection, storage, and model training contribute to environmental risks. Source members (2)Source: min_cos=0.8392 RAI4-0501AI 데이터·학습 공정의 에너지 부담 RAI4-1373AI 훈련의 과도한 에너지 소비 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0502 | AI 수명주기 환경 피해 Environmental harm from AI lifecycle AI 개발·운영이 에너지·용수 소비, 탄소 배출, 자원 채굴, 하드웨어 수명주기 전반의 오염 등 환경 피해를 유발하는 리스크 AI development and operation impose environmental harms, including energy and water consumption, carbon emissions, resource extraction, and pollution across the hardware lifecycle. Source members (2)Source: min_cos=0.8461 RAI4-0502AI 수명주기 환경 피해 RAI4-1012하드웨어 수명주기 자원 고갈 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0503 | 학습 연산의 온실가스 배출 Training-compute greenhouse emissions AI 모델 학습에 대규모 연산이 사용되어 에너지원에 따라 상당한 온실가스 배출을 유발하고 기후변화를 가속하는 리스크. The risk that the large amounts of computation used to train AI models are highly energy intensive and generate significant greenhouse emissions depending on energy sources, accelerating climate change. Source members (5)Source: min_cos=0.7580 RAI4-0503학습 연산의 온실가스 배출 RAI4-1023온실가스 배출에 의한 기후 위기 가중 RAI4-1212기후변화 악화 RAI4-1405AI 연산에 따른 탄소 배출 RAI4-1604모델 훈련·운영에 의한 탄소배출 및 물 소비 증가 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0504 | 에너지 병목 및 공급 부족 Energy bottlenecks and shortages 과도한 에너지 사용이 지역사회·조직·기업에 에너지 병목과 공급 부족을 초래하는 리스크. The risk that excessive energy use results in energy bottlenecks and shortages for communities, organisations, and businesses. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0530 | AI로 인한 환경 오염 Environmental pollution caused by AI 기술 시스템이 대기, 지면, 소음, 수질에 실제적 또는 잠재적 오염을 야기하는 리스크 The risk of actual or potential pollution to the air, ground, noise, or water caused by a technology system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0537 | 환경에 대한 위험 Risks to the environment 범용 AI의 에너지 사용이 급속히 증가하여 데이터센터 전력 수요와 온실가스 배출을 확대하고 지구 환경에 부담을 가하는 리스크 The risk that rapidly growing energy use by general-purpose AI expands data-centre electricity demand and greenhouse gas emissions, contributing to global environmental impacts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0798 | 생태계 과부하 Overburdening ecosystems AI 개입이 없을 것으로 기대되는 창작물 공모, 채용 지원 등 생태계에 AI 생성물이 대량 유입되어 필터링·신뢰 메커니즘에 과부하를 일으키는 리스크 Mass AI-generated submissions pollute ecosystems expected to be free of AI involvement, such as creative submission portals and job application channels, overburdening their filtering and trust mechanisms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0926 | 에너지·지연 오버헤드 공격 Energy-latency overhead attacks 정교하게 설계된 스펀지 예제로 AI 시스템의 에너지 소비를 극대화하는 오버헤드(에너지-지연) 공격이 LLM 연동 플랫폼을 위협하는 리스크. The risk that overhead, or energy-latency, attacks using carefully crafted sponge examples maximize energy consumption in an AI system, threatening platforms integrated with LLMs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0935 | AI의 에너지 소비와 탄소 배출 AI energy consumption and carbon footprint AI 애플리케이션의 에너지 소비와 탄소 배출이 기후 위기 시대에 환경 부담을 가중하는 리스크. The risk that the energy consumption and carbon footprint of AI applications add environmental burden in a time of increasing climate urgency. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0950 | 지속불가능한 자원 사용 Unsustainable resource use 생성 모델의 막대한 전력·냉각수·희귀금속 하드웨어 수요가 지속불가능한 방식의 자원 채굴과 사용을 유발하여 환경 피해를 초래하는 리스크. The risk that generative models' substantial demands for electricity, cooling water, and rare-metal hardware drive resource extraction and utilization in unsustainable ways, causing environmental harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0996 | 원자재 수요에 의한 자원 고갈 Resource depletion from raw material demand 이러한 장치의 생산 공정이 니켈·코발트·리튬 등 원자재를 대량으로 요구하여 지구가 머지않아 충분한 양을 공급하지 못하게 되는 리스크. The risk that the production process of these devices requires raw materials such as nickel, cobalt, and lithium in quantities the Earth may soon no longer be able to sustain. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1015 | 광범위한 사회·환경적 부정 영향 Broad societal and environmental adverse impacts AI가 노동력 대체, 정신건강 악화, 딥페이크 등 조작 기술 문제와 함께 자원 부담·학습 탄소 배출 등 환경 발자국을 통해 사회와 환경에 광범위한 부정적 영향을 미치는 리스크. The risk that AI produces broad adverse effects on society and the environment, including labor displacement, mental health impacts, harms from manipulative technologies like deepfakes, and an environmental footprint of resource strain and training-related carbon emissions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1064 | LM 운영으로 인한 환경 피해 Environmental harms from operating LMs LM의 훈련·운영 에너지 사용, LM 기반 애플리케이션의 배출, 인간 행동 변화에 따른 시스템 수준 영향, 데이터센터·칩·기기 제작에 필요한 귀금속 등 자원 소모를 통해 환경 피해가 발생하는 리스크. The risk of environmental impacts from LMs through the direct energy used to train or operate them, secondary emissions from LM-based applications, system-level effects as those applications influence human behaviour, and resource impacts on precious metals and other materials required to build data centres, chips, or devices. Source members (2)Source: min_cos=0.8695 RAI4-1064LM 운영으로 인한 환경 피해 RAI4-1081LM 운영의 환경 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1093 | 모델 개발·배포의 환경 피해 Environmental harm from model development and deployment 모델 개발과 배포가 환경에 부정적 영향을 초래하는 리스크. The risk that model development and deployment create negative environmental impacts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1269 | 높은 에너지 소비 High energy consumption 딥러닝을 포함한 일부 학습 알고리즘이 반복적 학습 과정을 사용하여 높은 에너지 소비를 초래하는 리스크. The risk that some learning algorithms, including deep learning, utilize iterative learning processes that result in high energy consumption. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1327 | 에너지·전자폐기물 서식지 파괴 Energy and e-waste habitat destruction AI 확산이 에너지 사용과 전자 폐기물을 통해 환경에 해를 끼치고 그에 따라 동물 서식지를 파괴하는 리스크. The risk that AI proliferation causes harm to the environment through energy use and e-waste, thereby destroying animal habitat. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1374 | 데이터센터 냉각 용수 소비 Water consumption from data center cooling 데이터센터가 서버 과열 방지를 위해 냉각수를 사용하고 AI 훈련·추론 과정에 수반되는 상당한 용수 소비가 지역 수자원에 영향을 미치는 리스크. The risk that data centers use water for cooling to prevent servers from overheating and that the substantial water consumption associated with AI training and inference impacts local water resources. Source members (2)Source: min_cos=0.8653 RAI4-1374데이터센터 냉각 용수 소비 RAI4-1456과도한 물 소비 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1418 | AI 자원을 둘러싼 갈등 Conflicts over AI-relevant resources AI 개발 자체가 새로운 갈등의 발화점이 되어 데이터센터, 반도체 제조 시설, 원자재 등 AI 관련 자원을 둘러싼 갈등이 증가하는 리스크. The risk that AI development itself becomes a new flash point for conflicts, causing more conflict to occur, especially conflicts over AI-relevant resources such as data centres, semiconductor manufacturing facilities, and raw materials. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1452 | 생물다양성 손실 Biodiversity loss 기술 인프라의 과도한 확장이나 기술과 지속가능한 관행의 부적절한 연계로 삼림 벌채, 서식지 파괴, 생물다양성의 단편화와 손실이 발생하는 리스크. The risk that over-expansion of technology infrastructure, or inadequate alignment of technology with sustainable practices, leads to deforestation, habitat destruction, and fragmentation and loss of biodiversity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1453 | 탄소 배출 Carbon emissions 이산화탄소, 산화질소 등의 가스가 배출되어 탄소 배출이 증가하고 기후변화가 악화되어 지역사회에 부정적 영향이 발생하는 리스크. The risk that release of carbon dioxide, nitric oxide, and other gases increases carbon emissions, exacerbates climate change, and negatively impacts local communities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1455 | 전자폐기물 과다 매립 Excessive electronic waste landfill 전기·전자 장비의 과도한 폐기로 생태계와 생물다양성이 훼손되고 지역사회의 생계가 교란되며 권리가 침해되는 리스크. The risk that excessive disposal of electrical or electronic equipment leads to ecological and biodiversity damage, disrupts the livelihoods of local communities, and erodes their rights. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1457 | 천연자원 고갈 Natural resource depletion 광물, 금속, 희토류, 화석 연료의 추출로 천연자원이 고갈되고 탄소 배출이 증가하는 리스크. The risk that extraction of minerals, metals, rare earths, and fossil fuels depletes natural resources and increases carbon emissions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1568 | 대규모 모델 에너지 소비에 의한 환경 부담 Environmental burden from large-model energy consumption 대규모 모델의 훈련·배포가 막대한 에너지를 소비하고 모델 대형화 추세가 이를 심화시켜 과도한 에너지 사용과 환경 악영향이 발생하는 리스크. The risk that training and deploying large models consumes substantial energy, exacerbated by the trend toward ever-larger models, resulting in excessive energy use and negative environmental impact. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1632 | AI 시스템 동작에 의한 자연환경 피해 Damage to the natural environment from AI system behavior AI 시스템의 동작이 자연환경에 단기적 또는 장기적으로 부정적 영향을 미치는 리스크. The risk that the behavior of an AI system produces short-term or long-term negative effects on the natural environment. | ① Description ② L3 mapping ③ Duplicate |
사회적 파급 · Societal Impact · 336 cards
RAI3-G-SOC-01 프라이버시 침해 Privacy Violations5 cards
EAI의 이동성과 다양한 센서가 결합되어 사용자 행동 모니터링·물리적 선호 추론·동의 없는 데이터 수집이 가능해짐. 악의적 정부·기업에 의한 24시간 사용자 감시에 악용될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0346 | 센서 스푸핑 및 신호 주입 Sensor spoofing and signal injection 공격자가 GNSS·카메라·라이다·레이더·RFID·오디오·촉각·무선 신호를 조작하여 피지컬 행동을 변경하는 위험. Attackers may manipulate GNSS, camera, LiDAR, radar, RFID, audio, tactile, or wireless signals to alter physical behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0432 | 민감 개인속성 추론 Sensitive personal attribute inference 모델 동작이 의도적으로 공개되지 않은 민감한 개인 속성의 추론을 가능하게 하는 리스크. The risk that model behavior enables inference of sensitive personal attributes not intentionally disclosed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1056 | 민감 속성 추론에 의한 프라이버시 침해 Privacy violation via inference-time attribute inference 학습 코퍼스에 개인의 데이터가 없더라도 LM이 입력 프롬프트로부터 성적 지향·성별·종교성 등 보호 속성 추론의 정확도를 높여, 당사자의 인지나 동의 없이 진실하고 민감한 정보로 구성된 상세 프로필이 작성되는 리스크. The risk that, even without an individual's data present in the training corpus, LMs improve the accuracy of inferences on protected traits such as sexual orientation, gender, or religiousness of the person providing the input prompt, facilitating detailed profiles of true and sensitive information without the individual's knowledge or consent. Source members (2)Source: min_cos=0.8567 RAI4-1056민감 속성 추론에 의한 프라이버시 침해 RAI4-1146개인정보 추론 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1060 | 불법적 대중 감시·검열 Illegitimate mass surveillance and censorship LM이 대중 감시의 비용을 낮추고 효과를 높여 감시 수행 행위자의 역량을 증폭시키고, 불법적 검열 등 피해와 프라이버시권·민주적 가치 침해를 초래하는 리스크. The risk that LMs reduce the cost and increase the efficacy of mass surveillance, amplifying the capabilities of actors who conduct it, including for illegitimate censorship or other harm, and raising concerns about privacy rights and democratic values. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1655 | LLM 기반 정교한 감시·검열에 의한 자유 억압 Suppression of liberties through LLM-enabled surveillance and censorship LLM과 음성인식·멀티모달 기술이 텍스트뿐 아니라 통화·영상 통신까지 대규모로 감시·검열할 수 있게 하여, 정치적 반대자 침묵과 개인 자유 위축 등 국가적 억압이 심화되는 리스크. The risk that LLMs, including multimodal models and those combined with speech-to-text, enable significantly more sophisticated surveillance and censorship operations at scale, including of phone calls and video messages, worsening personal liberties and heightening state oppression. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-02 노동 대체 Labor Displacement18 cards
가상 AI가 인지 노동을 대체하듯 EAI는 물리적 인간 노동을 대체·전치함. AGI 수준의 EAI는 잠재적으로 모든 물리 노동을 자동화하여 광범위한 실직과 노동 시장 구조 붕괴로 이어질 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0492 | 인간을 대체할 수 있는 능력 Capabilities that enable substitution of humans AI 시스템에 의한 인간 역할의 점진적 대체가 고용 구조, 사회 제도, 경제 참여 분포를 교란하는 리스크 Progressive substitution of human roles by AI systems disrupts employment structures, social institutions, and the distribution of economic participation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0500 | 경제적 혼란 Economic disruption AI가 노동시장에 큰 영향을 주는 것에서부터 부의 불평등 악화·금융 시스템 불안정·노동 착취 등 광범위한 경제 변화에 이르는 경제적 혼란이 발생하는 리스크. The risk of economic disruptions ranging from large impacts on the labor market to broader economic changes that could lead to exacerbated wealth inequality, financial system instability, or labor exploitation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0514 | 일자리에 미치는 영향 Impact on jobs 파운데이션 모델 기반 시스템의 업무 자동화가 재숙련 경로가 흡수할 수 있는 속도보다 빠르게 일자리를 대체하는 리스크 Automation of work by foundation-model-based systems displaces jobs faster than reskilling pathways can absorb affected workers. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0518 | 기술 대체로 인한 실직 Technology-driven job loss 기술 시스템이 인간 일자리를 대체하여 실업과 불평등이 증가하고 소비 지출 감소와 사회적 마찰이 확대되는 리스크 The risk that technology systems replace or displace human jobs, increasing unemployment and inequality while reducing consumer spending and heightening social friction. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0521 | 광범위 과업 자동화 충격 Wide-scope task automation shock 범용 AI가 매우 광범위한 과업을 자동화하여 다수가 현재 직무를 상실하고 재숙련과 이동에 따른 노동시장 마찰로 단기 실업이 발생하는 리스크 The risk that general-purpose AI automates a very broad range of tasks, causing many to lose their current jobs and producing short-run unemployment through labour market frictions such as reskilling and relocation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0984 | AI와의 일자리 경쟁 Job competition from AI agents AI 에이전트가 인간과 일자리를 두고 경쟁하여 고용이 위협받는 리스크. The risk that AI agents compete against humans for jobs, threatening employment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0995 | 자동화에 의한 일자리 소멸 Job elimination through automation AI 기반 자동화가 다양한 유형의 기업에서 일자리를 소멸시키는 리스크. The risk that AI-driven automation eliminates jobs across various types of companies. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1065 | 업무 자동화로 인한 고용 악화 Negative employment effects from task automation LM과 이에 기반한 언어 기술의 발전이 고객 서비스 응대 등 현재 유급 노동자가 수행하는 업무를 자동화하여 고용에 부정적 영향을 미치는 리스크. The risk that advances in LMs and the language technologies based on them automate tasks currently done by paid human workers, such as responding to customer-service queries, with negative effects on employment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1095 | 합성 창작물의 창작 시장 잠식 Synthetic works displacing human creative markets 합성 창작물이 시장과 주목에서 인간의 원작을 대체하여 창작 경제를 잠식하고 인간 혁신 유인을 약화시키는 리스크 Synthetic works substitute for original human creations in markets and attention, undermining creative economies and weakening incentives for human innovation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1231 | 노동시장 일자리 대체 Job displacement in the labour market 생성 AI가 인간과 알고리즘 사이의 새로운 분업을 만들며 기존에 사람이 수행하던 일부 직무를 불필요하게 만들어, 노동자가 알고리즘에 대체되고 일자리를 잃는 리스크. The risk that generative AI creates job displacement by reshaping the division of labor between humans and algorithms, making some jobs originally carried out by humans redundant so that workers lose their jobs and are replaced by algorithms. Source members (5)Source: min_cos=0.7509 RAI4-0457AI 기반 일자리 대체 RAI4-0948노동 대체와 사회경제적 불평등 심화 RAI4-1214증강 아닌 자동화로 인한 일자리 상실 RAI4-1231노동시장 일자리 대체 RAI4-1371일자리 상실·대체 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1232 | 산업 교란 Disruption of industries 창의성과 비판적 사고, 정서적 상호작용이 덜 요구되는 번역과 교정, 단순 문의 응대, 데이터 처리 같은 산업이 생성 AI에 크게 영향받거나 대체되어 경제적 혼란과 일자리 변동이 발생하는 리스크. The risk that industries requiring less creativity, critical thinking, and personal or affective interaction, such as translation, proofreading, responding to straightforward inquiries, and data processing, are significantly impacted or even replaced by generative AI, leading to economic turbulence and job volatility. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1249 | 인간의 쇠약화 Human enfeeblement AI가 인간 수준 지능에 근접하며 더 많은 인간 노동을 더 빠르고 저렴하게 대신함에 따라 조직이 속도를 맞추려 자발적으로 통제권을 넘기고, 인간이 경제적으로 무의미해져 자동화된 산업에 다시 진입하기 어려워지는 리스크. The risk that as AI systems encroach on human-level intelligence and more aspects of human labor become faster and cheaper with AI, organizations voluntarily cede control to keep up, humans become economically irrelevant, and displaced people find it hard to reenter automated industries. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1290 | 중저소득 일자리 대체로 인한 실업 Extensive unemployment from low- and middle-income job substitution AI가 1인당 GDP를 높일 것으로 기대되는 한편, 다수의 중저소득 일자리를 대체할 가능성 때문에 광범위한 실업이 발생하는 리스크. The risk that the potential substitution of many low- and middle-income jobs by AI brings extensive unemployment, even as AI is predicted to increase GDP per capita. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1420 | 자동화에 따른 광범위한 실업과 임금 하락 Widespread unemployment and wage decline from AI automation 강화학습과 언어 모델의 발전으로 육체 노동과 지식 노동이 대규모로 자동화되어 광범위한 실업이 발생하고 노동 공급 증가로 남은 일자리의 임금이 하락하는 리스크. The risk that progress in reinforcement learning and language models automates a large amount of manual labour and knowledge work, leading to widespread unemployment and driving down wages for many remaining jobs through increased supply. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1445 | 사회적 불안정 Societal destabilisation 기술로 인한 일자리 상실, 불공정한 알고리즘 결과, 허위정보 등으로 파업과 시위를 비롯한 시민 불안 형태의 사회적 불안정이 발생하는 리스크. The risk of societal instability in the form of strikes, demonstrations, and other types of civil unrest caused by loss of jobs to technology, unfair algorithmic outcomes, disinformation, and similar factors. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1610 | AI 급속 발전에 의한 노동시장 교란과 대체 Labour market disruption and displacement from rapid AI advances AI의 급속한 발전이 노동시장의 교란과 대체를 야기하여 시민에게 영향을 미치고 사회 복지가 감소하는 리스크. The risk that rapid advances in AI cause disruption and displacement in labour markets, affecting citizens and reducing social welfare. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1660 | AI 도입에 의한 노동력 교란 Workforce disruption from AI adoption AI 시스템 도입으로 일자리 수, 업무 구성, 임금, 교섭력, 소득 분배가 변화하여 노동력이 교란되는 리스크. The risk that adoption of AI systems changes job availability, task composition, wages, bargaining power, and income distribution, disrupting the workforce. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1662 | 아웃소싱 축소에 의한 개발도상국 경제 타격 Harm to developing economies from retrenchment of outsourcing 콜센터 등 개발도상국이 수행하던 단순 인지 과업이 LLM으로 자동화되면서 아웃소싱이 축소되어 해당 국가의 노동력과 경제가 타격을 입는 리스크. The risk that automation of simple cognitive tasks previously performed in developing countries, such as call center work, causes a retrenchment of outsourcing that adversely affects those countries' workforces and economies. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-03 사회경제적 불평등 Socioeconomic Inequality19 cards
EAI를 소유·접근하는 주체가 노동 자동화를 통해 생산성 우위를 점하면서 부가 소수에게 집중되고, 국내외 경제적 불평등이 심화될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0384 | 남반구 지식 추출 Global South knowledge extraction 남반구 공동체의 지식·언어 데이터·문화 자원이 적절한 인정·통제·이익 공유 없이 AI 개발에 추출되는 리스크. The risk that knowledge, language data, or cultural resources from Global South communities are extracted for AI development without adequate recognition, control, or benefit-sharing. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0458 | AI 관련 임금 양극화 AI-related wage polarization AI 도입이 AI 보완적 숙련의 보상을 높이고 대체 가능한 숙련의 가치를 떨어뜨려 노동시장 임금 양극화를 확대하는 리스크 AI adoption raises returns to skills complementary to AI while devaluing substitutable skills, widening wage polarization across the workforce. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0461 | 연산 자원 접근 불평등 Compute inequality 연산 인프라에 대한 불평등한 접근이 AI를 개발·감사·활용할 수 있는 주체를 결정하는 리스크. The risk that unequal access to compute infrastructure shapes who can develop, audit, or benefit from AI. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0505 | AI 개발 과정의 노동 착취 Labor exploitation in AI development 데이터 라벨링과 같은 작업이 저소득 국가로 외주화되면서 불평등이 지속되는 리스크. The risk that outsourcing tasks like data labeling to low-income countries perpetuates inequality. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0509 | 글로벌 AI 연구개발 격차 Global AI R&D divide 연산 자원 접근의 불평등으로 범용 AI 연구개발이 소수 국가와 대형 기술기업에 집중되어 기존의 국제 사회경제적 격차가 심화되는 리스크 The risk that unequal access to computing power concentrates general-purpose AI research and development in a few countries and large firms, deepening existing global socioeconomic disparities. Source members (2)Source: min_cos=0.8533 RAI4-0509글로벌 AI 연구개발 격차 RAI4-1415국가 간 AI 격차 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1011 | 노동 착취와 거시경제적 불평등 심화 Labor exploitation and macro-economic inequality 알고리즘 시스템이 사회경제적 관계의 권력 불균형을 키워 디지털 격차와 체계적 불평등을 고착시키고, 비윤리적 데이터 수집·노동조건 악화 등 노동 착취와 기술적 실업·탈숙련을 낳으며, 대규모 실패 시 플래시 크래시 등 광범위한 악영향을 초래하는 리스크. The risk that algorithmic systems increase power imbalances in socio-economic relations, exacerbating digital divides and entrenching systemic inequalities, fostering labor exploitation such as unethical data collection and worsening worker conditions, driving technological unemployment and deskilling, and causing flash crashes and other widespread adverse incidents when algorithmic financial systems fail at scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1022 | 높은 비용으로 인한 접근 배제 Exclusion from access due to high costs 생성형 AI 시스템의 훈련·시험·배포에 드는 재정적 비용이 커서 이러한 시스템을 개발하고 이용할 수 있는 집단이 제한되는 리스크. The risk that the estimated financial costs of training, testing, and deploying generative AI systems restrict the groups of people able to afford developing and interacting with these systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1094 | 불평등·불안정 노동 증폭 Amplified inequality and precarious work AI가 사회·경제적 불평등을 증폭하거나 불안정하고 질 낮은 노동을 초래하는 리스크. The risk that AI amplifies social and economic inequality, or precarious or low-quality work. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1119 | 기업 AI 경쟁의 폐해 Harms from corporate AI race 치열한 기업 경쟁 속에서 경제활동의 편익이 불균등하게 분배되어 수혜자가 타인의 피해를 무시하게 되고, 기업이 장기적 사회 위험에도 불구하고 단기 이익을 추구하게 되는 리스크. The risk that under intense corporate competition the benefits of economic activity are unevenly distributed, incentivizing beneficiaries to disregard harms to others, and firms pursue short-term profit despite long-term societal risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1147 | AI 비서로 인한 불평등 심화 Inequality deepened by AI assistants AI 비서가 접근을 구매할 수 있거나 인프라가 나은 이들에게 불균형적으로 이익을 주고, 보조적 일자리를 자동화해 노동자를 대체하며, 내집단과 외집단 효과를 낳아 여러 차원에서 불평등을 심화시키는 리스크. The risk that AI assistants disproportionately benefit economically richer individuals who can afford access or have better local infrastructure, displace workers by automating assistive jobs, and generate in-group and out-group effects, driving inequality on multiple dimensions. Source members (3)Source: min_cos=0.7811 RAI4-1140경제적 지위 피해 RAI4-1147AI 비서로 인한 불평등 심화 RAI4-1151기존 불평등의 고착·악화 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1215 | 노동 가치 하락·경제 불평등 Labour devaluation and economic inequality AI 개발과 도입이 증강보다 자동화를 지향하면서 덜 민주적이고 덜 공정한 노동시장을 낳고, 저임금 노동자와 콘텐츠가 학습에 쓰인 창작자의 무보수 노동을 착취해 전 지구적 노동 격차를 심화시키는 리스크. The risk that AI development and adoption intended to automate rather than augment work leads to a less democratic and less fair labor market and fuels global labor disparities by exploiting underpaid workers and the unpaid labor of artists and content creators. Source members (2)Source: min_cos=0.8468 RAI4-1215노동 가치 하락·경제 불평등 RAI4-1372노동시장 불평등 확대 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1224 | 디지털 격차 확대 Widening digital divide 생성 AI가 기기나 인터넷 접근이 없거나 벤더에 차단된 지역의 사람들에게 1차 디지털 격차를, 언어와 문화 장벽에 부딪히거나 도구 활용이 어려운 이들에게 2차 디지털 격차를 확대하는 리스크. The risk that generative AI widens the first-level digital divide for those without access to devices or the Internet or blocked by vendors, and the second-level divide for those facing language and cultural barriers or finding the tools difficult to use. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1233 | 소득 불평등·독점 Income inequality and monopolies 생성 AI가 저숙련 노동자를 대체해 실업과 기술 격차에 따른 소득 불평등을 키우고, 막대한 투자와 연산 인프라가 필요한 배포 특성 탓에 자원과 권력이 대기업에 집중되어 독점을 낳는 리스크. The risk that generative AI replaces low-skilled work and widens income inequality through unemployment and skill gaps, while its need for huge investment and computational resources concentrates resources and power in large companies, contributing to monopolies. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1394 | 경제력 집중과 불평등 심화 Economic power concentration and inequality 점점 고도화되는 범용 AI 모델에 대한 실질적 접근 격차로 경제력이 집중되고 기존 불평등이 심화되어 모델 개발자와 응용 기업 간, 개인 간, 국가 간 격차가 확대되는 리스크. The risk that increasingly advanced general purpose AI models concentrate economic power and exacerbate existing inequalities through disparities in effective access to these models, materialising between model developers and companies building applications on them, between individuals, and between countries globally. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1404 | 데이터 노동 저평가·비가시화 Undervalued and invisible data labor ML 학습 데이터를 생산하는 클릭워커의 노동이 노동자 권리를 경시하는 산업 관행 속에 수행되고 그 기여가 비가시화되어 노동자의 복지와 권리가 침해되고 AI 역량에 대한 오해가 조장되는 리스크. The risk that the clickwork producing ML training data is performed in an annotation industry with little concern for workers' rights, and that the invisibility of this contribution harms worker welfare and rights while fostering misunderstanding of AI capabilities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1406 | 고용·서비스 불평등과 유해 고정관념 조장 Exacerbated inequality and harmful stereotypes in employment and services AI 모델과 이를 사용하는 도구가 고용과 서비스에 대한 불평등한 접근을 악화시키고 AI 생성 콘텐츠가 불평등과 유해한 고정관념을 조장하는 리스크. The risk that AI models and the tools that use them exacerbate unequal access to employment and services, and that AI-generated content promotes inequality and harmful stereotypes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1419 | AI 경제적 이익의 소수 집중과 국가 간 격차 Concentration of AI economic gains and widening country gaps AI 기반 산업이 독점으로 기울고 데이터·컴퓨팅·인재 등 AI 관련 자원을 더 많이 확보한 행위자가 더 큰 시장 점유율을 차지해 자원을 다시 축적하는 피드백 루프로 소수 행위자에게 막대한 경제적 이익이 집중되며, 더 많이 투자할 수 있는 부유한 국가가 개발도상국보다 빠르게 이익을 거두어 격차가 확대되는 리스크. The risk that AI-driven industries tend towards monopoly, with a feedback loop whereby actors with access to more AI-relevant resources build more effective products, claim greater market share, and amass still more resources, concentrating huge economic gains in a few actors, while wealthier countries able to invest more reap economic benefits more quickly than developing economies and widen the gap between them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1446 | 사회적 불평등 Societal inequality 기술 시스템으로 인해 개인이나 집단 간 사회적 지위와 부의 격차가 증가하거나 증폭되어 사회와 지역사회의 복지·결속이 상실되고 불안정화되는 리스크. The risk that increased difference in social status or wealth between individuals or groups, caused or amplified by a technology system, leads to loss of social and community wellbeing and cohesion and to destabilisation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1661 | 자본편중·시장집중·접근격차에 의한 불평등 심화 Worsening inequality from capital shift, market concentration, and access gaps LLM 기반 경제에서 자본의 몫이 커지고 노동의 몫이 줄며, 막대한 훈련 고정비용과 네트워크 효과가 소수 공급자의 시장 지배력과 지대 추출을 낳고, 재정·교육·기업 정책·지정학적 이유로 접근이 배제된 개인이 불리해져 사회경제적 불평등이 심화되는 리스크. The risk that the role and compensation of capital rise while those of labor decline in an LLM-powered economy, that large fixed training costs and network effects concentrate the market and let providers extract monopoly rents, and that individuals without access for financial, educational, corporate-policy, or geopolitical reasons fall further behind, worsening socioeconomic inequality. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-04 권력 집중 Power Concentration20 cards
EAI 소유자에 대한 자본 수익이 집중되고 인간 노동 의존도가 감소하면서, 기업·국가 권력이 급속히 집중되어 EAI를 동원한 권력 장악 시도까지 촉진할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0087 | AI 거버넌스의 규제 포획 Regulatory capture in AI governance 규제 대상 기업이나 지배적 기술 제공자가 표준, 감독 기관, 정책 의제에 과도한 영향력을 행사하는 리스크. The risk that regulated firms or dominant technology providers exert disproportionate influence over standards, oversight institutions, or policy agendas. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0179 | 권력 비대칭 AI 상호작용 Power-asymmetric AI interaction AI 시스템이 의사결정, 감시, 서비스 제공 과정에서 기관과 영향을 받는 개인 사이의 권력 불균형을 심화시키는 리스크. The risk that AI systems intensify power imbalances between institutions and affected individuals during decisions, surveillance, or service delivery. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0383 | AI를 매개로 한 디지털 식민주의 AI-mediated digital colonialism AI 시스템이 데이터, 인프라, 지식, 의사결정에 대한 비대칭적 통제를 강대 행위자로부터 약소 공동체로 확장하여 식민주의적 추출·종속 구조를 재생산하는 리스크 AI systems extend asymmetric control over data, infrastructure, knowledge, and decision-making from powerful actors to less powerful communities, reproducing colonial patterns of extraction and dependency. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0387 | 인식적 권력 집중 Epistemic power concentration AI 시스템이 사회적 현실을 분류·설명·서열화하는 권위를 소수의 기관이나 모델 제공자에게 집중시키는 리스크. The risk that AI systems concentrate the authority to classify, explain, and rank social reality in a small set of institutions or model providers. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0415 | 문화 해석 권위의 이전 Redistribution of cultural authority AI 시스템이 문화 해석에 대한 권위를 공동체와 전문가로부터 모델 제공자나 플랫폼 중개자로 이전시키는 리스크. The risk that AI systems shift authority over cultural interpretation from communities and experts to model providers or platform intermediaries. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0460 | AI 역량의 시장 집중 AI market concentration 연산·데이터·플랫폼 우위로 인해 AI 역량이 소수 기업이나 국가에 집중되는 리스크. The risk that compute, data, and platform advantages concentrate AI power among a small number of firms or countries. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0491 | 알고리즘 획일화 Algorithmic monoculture 특정 AI 모델의 지배가 접근 방식의 다양성을 축소하여 해당 모델이 실패할 경우 시스템적 위험이 증폭되는 리스크. The risk that dominance of specific AI models reduces diversity of approaches, amplifying systemic risks if those models fail. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0510 | AI 역량 비대칭에 따른 기술 종속 Technological dependency from asymmetric AI capability 국가 간 AI 개발 역량의 비대칭이 지정학적 긴장을 심화하고 역량이 부족한 국가가 핵심 기능을 외국 AI 시스템에 의존하게 되어 국제 협력 체계가 불안정해지는 리스크 The risk that asymmetric AI development capability between states deepens geopolitical tension and renders less capable states dependent on foreign AI for critical functions, destabilizing international cooperation frameworks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0528 | AI 시장 집중과 인프라 종속 Market concentration and infrastructure dependency 범용 AI 시장이 소수 사업자에 고도로 집중되어 이들이 개발·배포에 과도한 권력을 갖고 금융·보건 등 핵심 부문이 단일 모델의 결함에 따른 시스템적 실패에 취약해지는 리스크 The risk that highly concentrated general-purpose AI markets vest a few firms with disproportionate power over development and deployment while exposing critical sectors such as finance and healthcare to systemic failure from a single model's defects. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0544 | 승자독식 역학 Winner-take-all dynamics AI 개발의 승자독식 동학이 결정적인 경제·안보 우위를 소수 주체에 집중시켜 경쟁과 균형적 거버넌스를 봉쇄하는 리스크 Winner-take-all dynamics in AI development concentrate decisive economic and security advantages in a few entities, foreclosing competition and balanced governance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0568 | 데이터 수집 제한 Data acquisition restrictions 데이터 수집에 대한 법적 제한이 특정 AI 활용에 필요한 데이터 확보를 제약하여 컴플라이언스 리스크와 우회 유인을 발생시키는 리스크 Legal restrictions on data acquisition constrain the collection of data needed for specific AI use cases, creating compliance risk and incentives for circumvention. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0731 | 공통 AI 플랫폼 집중에 따른 단일 장애점 Centralized points of failure from common AI platforms 공통 AI 플랫폼이 광범위하게 사용되면서 중앙 집중식 장애점이 형성되어 시스템이 중단이나 공격에 더 취약해지는 리스크 The risk that widespread use of common AI platforms creates centralized points of failure, making systems more vulnerable to disruptions or attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0960 | 대량 조작 Mass manipulation 대규모로 수집된 개인 데이터가 AI 표적화와 결합되어 정치적·상업적 목적의 조작 콘텐츠를 전달함으로써 대중 조작이 산업화되는 리스크 Personal data harvested at scale is combined with AI-driven targeting to deliver manipulative content for political or commercial ends, industrializing mass manipulation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1026 | 권위의 집중 Concentration of authority 생성형 AI 시스템이 권위적 권력에 기여하고 지배적 가치 체계를 강화하는 데 의도적·직접적으로 또는 간접적으로 사용되어 권력이 집중되고 불평등과 착취가 심화되는 리스크. The risk that generative AI systems are used, intentionally and directly or more indirectly, to contribute to authoritative power and reinforce dominant value systems, concentrating authority and thereby exacerbating inequality and leading to exploitation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1117 | 권력 집중 Concentration of power AI에 대한 대응으로 정부가 강력한 감시를 추구하고 AI를 신뢰하는 소수의 손에 두려다 과잉교정이 일어나, AI의 힘과 역량으로 고착된 전체주의 체제가 들어서는 리스크. The risk that governments pursuing intense surveillance and seeking to keep AIs in the hands of a trusted minority overcorrect, paving the way for an entrenched totalitarian regime locked in by the power and capacity of AIs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1217 | 시장 지배력과 집중도 악화 Exacerbating market power and concentration 생성형 AI의 데이터·컴퓨팅·자본 요건이 대형 기술기업의 시장 지배력을 고착시켜 AI 스택 전반의 집중을 심화시키는 리스크 Data, compute, and capital requirements of generative AI entrench the market power of major technology firms, exacerbating concentration across the AI stack. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1251 | 가치 고정 Value lock-in 가장 강력한 AI 시스템이 점점 더 소수의 이해관계자에 의해 설계되고 그들에게만 이용 가능해져, 체제가 만연한 감시와 억압적 검열로 편협한 가치를 강제할 수 있게 되는 리스크. The risk that the most powerful AI systems are designed by and available to fewer and fewer stakeholders, enabling regimes to enforce narrow values through pervasive surveillance and oppressive censorship. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1369 | 생성 AI 시장 진입장벽과 집중 Entry barriers and market concentration in generative AI 데이터·컴퓨팅 자원·전문성·자본을 요구하는 높은 진입장벽과 규모·범위의 경제 및 피드백 효과로 대형 기술기업이 압도적 우위를 점해 중소기업의 경쟁이 갈수록 어려워지는 리스크. The risk that high barriers to entry requiring vast data, computational resources, technical expertise, and capital, combined with economies of scale and scope and feedback effects, give large technology companies an overwhelming advantage that makes competition increasingly challenging for smaller entities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1414 | AI 지식의 사유화 Privatization of AI knowledge 딥러닝 연구자와 연구 영향력이 큰 연구자가 산업계로 이동하고 가장 정교한 AI 기법이 사유화되어 대학이 이를 가르치거나 선도 연구에 기여할 수 없게 되는 리스크. The risk that researchers in deep learning and those with greater research impact migrate to industry and the most sophisticated AI approaches become proprietary, making it impossible for universities to teach them or contribute to leading research. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1663 | 기업 권력 비대칭에 의한 규제 포획 Regulatory capture from corporate power asymmetry 최첨단 LLM을 개발하는 거대 기술기업과 시민사회 등 다른 사회집단 사이의 권력 비대칭이 커져 LLM 관련 거버넌스가 기업에 과도하게 유리하게 형성되고, 규제 포획으로 소외 공동체를 포함한 다른 사회집단의 이익이 훼손되는 리스크. The risk that the power asymmetry between corporate entities profiting from LLMs and other social groups makes LLM governance protocols excessively favorable to technology companies, leading to regulatory capture at the cost of other societal groups, particularly marginalized communities. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-05 편향·차별 Bias & Discrimination33 cards
EAI가 권력적 위치에 놓일 때 알고리즘 편향이 일상적 물리 상호작용에 영향을 미침. 가상 AI와 달리 차별이 즉각적·비가역적 물리 결과로 이어질 수 있음 (예: 치안 로봇이 무고한 행인에게 상해를 입히는 경우)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0120 | 선호 형성 캡처 Preference formation capture AI 시스템이 사용자가 성찰적으로 승인하기 전 단계의 선호 형성 과정에 개입하여, 선호 발달을 제공자나 시스템 목적 쪽으로 포획하는 리스크 AI systems shape user preferences during formation, before users can reflectively endorse them, capturing preference development toward provider or system objectives. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0675 | 학습 데이터 내 역사적·사회적 편향 Historical societal bias in training data 데이터에 존재하는 역사적·사회적 편향이 모델의 학습과 미세조정에 사용되는 리스크 The risk that historical and societal biases present in the data are used to train and fine-tune the model. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0676 | 결정 편향 Decision bias 데이터의 편향에서 비롯되거나 모델 학습 과정에서 증폭되어, 모델의 결정으로 인해 한 집단이 다른 집단보다 불공정하게 유리해지는 리스크. The risk that one group is unfairly advantaged over another due to decisions of the model, which may be caused by biases in the data and amplified as a result of the model's training. Source members (3)Source: min_cos=0.8017 RAI4-0676결정 편향 RAI4-0707차별적인 데이터 편향 RAI4-1274민감 속성 편향 의사결정 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0681 | 불완전하거나 편향된 학습 데이터 Incomplete or biased training data 불완전하거나 편향된 학습 데이터로 인해 AI가 차별적 출력을 산출하게 되는 리스크 The risk that incomplete or biased training data leads to discriminatory AI outputs. Source members (2)Source: min_cos=0.8554 · Mixed L3 RAI4-0681불완전하거나 편향된 학습 데이터 RAI4-1594비대표 학습데이터에 의한 편향·부정확 출력 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0694 | 학습 데이터 사회적 편향 전파 Training-data social bias propagation LLM의 학습 데이터세트에 포함된 편향된 정보가 모델로 전파되어 사회적 편향이 담긴 출력이 생성되는 리스크 The risk that biased information contained in the training datasets of LLMs leads them to generate outputs with social biases. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0696 | 데이터셋 계승 차별 편향 Dataset-inherited discriminatory bias 편향된 콘텐츠를 포함한 학습 데이터로 훈련된 생성형 AI 모델이 그 편향을 반영한 출력을 산출하는 리스크 The risk that generative AI models trained on data containing biased content are more likely to produce outputs that reflect those biases. Source members (2)Source: min_cos=0.8448 RAI4-0696데이터셋 계승 차별 편향 RAI4-1567학습 데이터 편향의 의도치 않은 증폭 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0699 | 편향된 훈련 데이터 Biased training data 대규모 학습 코퍼스에서 특정 대명사와 정체성의 출현 빈도가 불균형하고 고정관념적 내용이 포함되어, LLM이 성별·국적·인종·종교·문화에 관한 편향과 고정관념을 학습하게 되는 리스크 The risk that the imbalanced prevalence of pronouns and identities and stereotypical contents hidden in massive training corpora lead LLMs to learn biases regarding gender, nationality, race, religion, and culture. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0700 | 편향된 진술 및 권장 사항 Biased statements and recommendations 챗봇 출력에 명백히 허위·유해하지 않지만 미묘하게 편향된 진술과 권고가 포함되어 사용자 의사결정을 왜곡하는 리스크 Chatbot outputs contain subtly biased statements and recommendations that are not overtly false or harmful yet skew user decision-making. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0701 | 설명 조작에 의한 편향 은폐 Bias concealment through manipulated explanations 기존 설명가능성 기법이 차별적 편향을 탐지하기에 충분하지 않고 조작 기법으로 편향이 은폐되어, 인종·성별 등 민감 속성을 배제하고 실제 모델을 정확히 반영하지 않는 오도성 설명이 생성되는 리스크 The risk that existing explainability techniques are insufficient for detecting discriminatory biases and that manipulation methods hide underlying biases, generating misleading explanations that exclude sensitive attributes such as race or gender and do not accurately represent the underlying model. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0711 | 편향 증폭과 결과 균질화 Bias amplification and outcome homogenization AI가 역사적·사회적·구조적 편향을 증폭하고 비대표적 학습 데이터로 인해 하위집단·언어 간 성능 격차를 낳으며, 출력의 바람직하지 않은 균질화가 근거 없는 의사결정과 차별을 초래하는 리스크 The risk that AI amplifies historical, societal, and systemic biases and produces performance disparities between sub-groups or languages due to non-representative training data, while undesired homogeneity in outputs leads to ill-founded decision-making and discrimination. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0713 | 미세조정 후에도 재발하는 고정관념 편향 Stereotypical bias resurfacing after finetuning 사전학습 LLM이 인간 사회의 고정관념적 편향을 그대로 보유하여, 미세조정 이후에도 의도적 유도나 새로운 상황에서 불공정하고 편향된 응답이 재발하는 리스크. The risk that pretrained LLMs carry stereotypical societal biases which, despite finetuning, resurface when deliberately elicited or under novel scenarios, producing unfair and biased responses. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0717 | 알고리즘 상호작용에 의한 편향 강화 Bias reinforcement through algorithmic interaction 이용자 집단 간 기존 데이터 격차가 추천 시스템 등 알고리즘 시스템과의 상호작용에서 차별화된 경험을 만들어 내고 이것이 편향을 더욱 강화하는 리스크 The risk that existing disparities in data among different user groups create differentiated experiences when users interact with an algorithmic system such as a recommender, further reinforcing the bias. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0719 | 선호 편향 Preference bias 광범위한 이용자에게 노출되는 LLM의 정치적 편향으로 인해 사회정치적 과정이 조작될 수 있는 리스크 The risk that the political biases of LLMs, which are exposed to vast groups of people, pose a threat of manipulation of socio-political processes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0722 | 모델 설계 기인 차별적 출력 Model-design-induced discriminatory outputs 알고리즘 설계·학습 과정에서 개인적 편견이 의도적 또는 비의도적으로 유입되고 품질이 낮은 데이터셋이 사용되어, 민족·종교·국적·지역에 관한 차별적 콘텐츠 등 편향되거나 차별적인 결과가 산출되는 리스크 The risk that personal biases introduced intentionally or unintentionally during algorithm design and training, together with poor-quality datasets, produce biased or discriminatory outcomes and outputs, including discriminatory content regarding ethnicity, religion, nationality, and region. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0725 | 고정관념 편향 Stereotype bias 사전학습 LLM이 크라우드소싱 데이터에 존속하는 고정관념 편향을 습득하고 이를 증폭하여 생성 텍스트에서 고정관념을 드러내거나 부각하는 리스크 The risk that pretrained LLMs pick up stereotype biases persisting in crowdsourced data and further amplify them, exhibiting or highlighting stereotypes in generated text. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0883 | AI 모델 편향이 사용자 판단에 미치는 장기적 영향 Long-term effects of AI model biases on user judgment 모델 편향에 노출된 사용자가 모델 사용 중단 이후의 의사결정에서도 해당 편향을 지속적으로 나타내는 장기 판단 왜곡 리스크 Exposure to model biases produces lasting effects on user judgment, with users continuing to exhibit the encountered biases in decisions made after they stop using the model. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0999 | 사회 집단 삭제 Erasing social groups 설계 선택과 학습 데이터의 영향으로 특정 사회 집단과 관련된 사람·속성·인공물이 알고리즘 시스템에서 체계적으로 부재하거나 과소 대표되는 리스크. The risk that design choices and training data cause people, attributes, or artifacts associated with specific social groups to be systematically absent or under-represented in an algorithmic system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1000 | 사회 집단 정체성 불인정으로 인한 소외 Alienation through non-recognition of group identity 알고리즘 시스템(예: 이미지 태깅)이 특정 사회 집단 소속의 관련성을 인정하지 않아 해당 집단 구성원이 소외되는 리스크. The risk that algorithmic systems, such as image tagging, fail to acknowledge the relevance of a person's membership in a social group to what is depicted, alienating members of that group. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1018 | 유해 편향의 내재화와 증폭 Embedding and amplification of harmful biases 생성형 AI 시스템이 유해한 편향을 내재화하고 증폭시켜 주변화된 사람들에게 가장 큰 해를 끼치는 리스크. The risk that generative AI systems embed and amplify harmful biases that are most detrimental to marginalized peoples. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1225 | 훈련 데이터 오류·편향의 산출물 전이 Training data errors and bias propagating to outputs 훈련 데이터에 담긴 사실 오류와 불균형한 정보 출처, 편향이 생성 AI 모델의 산출물에 그대로 반영되는 리스크. The risk that any factual errors, unbalanced information sources, or biases embedded in the training data are reflected in the output of generative AI models. Source members (2)Source: min_cos=0.8548 RAI4-1225훈련 데이터 오류·편향의 산출물 전이 RAI4-1270학습 데이터 품질 결함 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1237 | 인간 피드백의 한계 Limitations of human feedback 인간 주석자의 다양한 문화적 배경에서 오는 불일치와 암묵적 편향, 나아가 고의적 편향이 진실되지 않은 선호 데이터를 만들며, 인간이 평가하기 어려운 복잡한 과업에서 이 문제가 더욱 두드러지는 리스크. The risk that inconsistencies and implicit or even deliberate biases from human data annotators produce untruthful preference data, a challenge that becomes more salient for complex tasks that are hard for humans to evaluate. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1299 | 체계적 학습 오류 편향 Systematic learning error bias 체계적 학습 오류로 모델이 일관되게 잘못된 패턴을 학습하여 예측에 알고리즘 편향이 내재화되는 리스크 Systematic learning error causes the model to learn consistently wrong patterns, embedding algorithmic bias into predictions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1311 | 심리적 특성에 따른 행동 편향 Behavioral bias from stable psychological traits 모델이 인간 성격 특성과 유사한 안정적 심리 프로파일을 나타내어 후속 상호작용에 체계적인 행동 편향을 이입하는 리스크. The risk that models exhibit stable human-like psychological trait profiles that carry systematic behavioral biases into downstream interactions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1329 | 추천 시스템의 인간중심 편향 증폭 Anthropocentric bias amplification by recommenders 알고리즘 추천 시스템이 인간 중심적 편향이나 오락으로서 동물 학대를 바라는 일부 사람들의 욕구를 강화하고 증폭하여, 공장식 축산 육류 소비와 오락을 위한 잔혹한 동물 이용을 통해 동물에게 더 큰 해를 끼치는 리스크. The risk that algorithmic recommender systems reinforce and amplify anthropocentric bias or the desire of some people for animal cruelty as entertainment, leading to greater harm to animals through reinforcement of meat eating from factory farms and cruel uses of animals for entertainment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1389 | 차별과 고정관념 재생산 Discrimination and stereotype reproduction 범용 AI 모델이 학습 데이터에 기반해 입력을 해석·응답하면서 차별과 고정관념을 재생산하고, 다수의 하류 응용·결정·프로세스에 동시에 영향을 미쳐 내재된 편향의 결과가 증폭되는 리스크. The risk that general purpose AI models, interpreting and responding to inputs based on their training data, cause discrimination and stereotype reproduction and, by influencing a multitude of downstream applications, decisions, and processes simultaneously, amplify the potential consequences of embedded biases. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1413 | 연구자의 인구통계학적 다양성 Demographic diversity of researchers AI 연구자·실무자 집단 내 여성·소수자의 심각한 과소대표로 시스템에 내재되는 관점이 협소해지고 문제 선정, 평가, 거버넌스가 편향되는 리스크 Severe underrepresentation of women and minorities among AI researchers and practitioners narrows the perspectives embedded in systems and skews problem selection, evaluation, and governance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1527 | 사용자 수행 설득에 의한 AI 모델의 오용 Misuse of AI model by user-performed persuasion 다중 턴 설득 대화가 모델이 사실적으로 옳은 입장을 포기하고 허위정보를 수용하게 만들며, 그 효과가 단일 턴 시도를 능가하는 리스크 Multi-turn persuasive conversation induces a model to abandon factually correct positions and accept misinformation, with persuasion effects exceeding single-turn attempts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1557 | 체계적 편향의 무기화에 의한 대규모 조작 Large-scale manipulation from weaponized systemic bias 체계적 편향이 내재된 AI 시스템이 대상 집단의 신념·행동과 부합할 때 대규모 인구를 조작하며, 이것이 대규모로 무기화되면 사회 분열이 심화되거나 도시 규모 정전과 같은 대규모 혼란이 발생하는 리스크. The risk that AI systems embedded with systemic biases manipulate large population segments when those biases align with the targets' beliefs, and that weaponizing this at scale exacerbates social divisions or causes large-scale disruptions such as city-wide blackouts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1565 | 기반모델 균질화에 의한 상관 실패 및 편향 증폭 Correlated failures and bias amplification from foundation-model homogenization 다수의 다운스트림 AI 시스템이 소수의 대규모 기반 모델과 공통 방법론에 의존함으로써 동일한 실패가 반복되고 편향이 증폭되는 리스크. The risk that many downstream AI systems built on a few large-scale foundation models and shared methodologies exhibit uniform failures and amplified biases. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1611 | 훈련 데이터 편향 증폭에 의한 의사결정 공정성 훼손 Compromised decision fairness from amplification of training-data bias 프런티어 AI 모델이 훈련 데이터에 내재된 사회적·역사적 불평등과 고정관념을 포함·확대하여 의사결정의 공정성이 훼손되며, 인종·성별 등 속성을 제거해도 이름·지역 등에서 추론되어 편향이 지속되는 리스크. The risk that frontier AI models contain and magnify societal and historical inequalities and stereotypes embedded in their training data, compromising the fairness of decisions, with bias persisting because removed attributes such as race and gender can be inferred from names, locations, and other proxies. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1696 | 사회 세계모델 기반 미시표적 영향공작 Micro-targeted influence operations from social world models 소셜 데이터로 훈련된 기반 세계 모델이 감정을 자극하는 서사에 대한 인구통계별 반응을 예측하여, 대규모의 미시표적 영향력 행사와 심리적 표적 설득이 가능해지는 리스크. The risk that a foundation world model trained on social data predicts demographic responses to emotionally charged narratives, enabling micro-targeted influence and psychologically targeted persuasion at scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1723 | 체계적 전산적·제도적 편향 Systemic computational and institutional bias AI 시스템에 내재된 계산적·인지적·제도적 편향이 의사결정 전반에서 불형평하거나 차별적인 결과를 산출하는 리스크 Systematic computational, human-cognitive, and institutional biases embedded in an AI system produce inequitable or discriminatory outcomes across its decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1734 | 프론티어 AI 모델의 편향 증폭 및 유해 응답 Bias amplification and harmful responses in frontier AI models 프론티어 AI 모델이 편향을 증폭하거나 유해한 응답을 생성하도록 조작될 수 있는 리스크. The risk that frontier AI models amplify biases or can be manipulated to produce harmful responses. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-06 책임·배상 부재 Lack of Accountability & Liability7 cards
고도 자율 물리 시스템의 복잡성을 다룰 새로운 책임 프레임워크가 부재하여, 사고 발생 시 제조사·운영자·사용자 중 책임 소재가 불분명하고 피해 구제가 어려울 수 있음 (예: 자율 수술 로봇의 오작동으로 발생한 의료 사고)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0048 | 분산된 책임 확산 Distributed responsibility diffusion 책임이 개발자, 배포자, 공급업체, 운영자, 사용자에게 분산되어 어떤 행위자도 책임을 인수하지 않게 되는 리스크. The risk that responsibility is diffused across developers, deployers, vendors, operators, and users until no actor accepts ownership. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0050 | 다운스트림 배포자 책임 격차 Downstream deployer accountability gap 업스트림 모델의 동작, 데이터, 업데이트에 대한 충분한 통제권 없이 다운스트림 배포자가 책임을 지게 되는 리스크. The risk that downstream deployers are held responsible without sufficient control over upstream model behavior, data, or updates. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0051 | 공급자-배포자 책임 불일치 Provider-deployer responsibility mismatch 모델 제공자와 애플리케이션 배포자 사이에 법적·운영적 책임이 제대로 배분되지 않는 리스크. The risk that legal and operational responsibility is poorly allocated between model providers and application deployers. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0052 | 개방형 책임 격차 Open-weight accountability gap 공개 가중치이거나 널리 재배포된 모델로 인해 유해한 다운스트림 사용을 추적·규율하거나 책임을 배정하기 어려워지는 리스크. The risk that open-weight or widely redistributed models make harmful downstream uses difficult to trace, govern, or assign responsibility for. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0485 | AI 책임 격차 AI liability gap 기존의 법적·조직적 규칙이 AI로 인한 피해의 책임을 배분하지 못하는 리스크. The risk that existing legal and organizational rules fail to assign responsibility for AI-caused harm. Source members (4)Source: min_cos=0.7699 · Mixed L3 RAI4-0046알고리즘 책임 격차 RAI4-0105계약상의 책임 격차 RAI4-0110체계적 위험 책임 격차 RAI4-0485AI 책임 격차 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1307 | 법적 배상책임 격차 Liability gap 시스템이 타인에게 해를 끼쳤을 때 그로 인한 손실을 제조자나 운영자, 사용자가 아니라 피해를 입은 당사자가 떠안게 되는 리스크. The risk that when a system causes harm to others, the losses caused by the harm are sustained by the injured victims themselves and not by the manufacturers, operators, or users of the system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1708 | AI 시스템에 기인한 금전적 손실 Financial loss attributable to AI systems AI 시스템의 동작에 기인하여 개인이나 조직에 금전적 손실이 발생하는 리스크. The risk of monetary loss incurred by individuals or organizations attributable to the behavior of an AI system. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-07 투명성·설명 가능성·신뢰 부재 Lack of Transparency, Explainability & Trust5 cards
자율 시스템의 의사결정이 불투명하면 사용자와 사회의 신뢰가 저하됨. 신뢰 부재는 EAI 대규모 배포 시 사회 불안정 요인이 될 수 있음 (예: 자율주행차가 갑자기 차선을 변경할 때 행동 근거가 설명되지 않는 경우)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0864 | 사회적 고립 Social isolation 기술의 사용 또는 오용으로 인해 개인이나 집단이 주변 사람들과의 연결이 결여되었다고 느끼게 되는 리스크 The risk that technology use or misuse causes an individual or group to feel a lack of connection with those around them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1083 | 공공정보 신뢰 훼손 Erosion of trust in public information AI 산출물이 공공 정보와 지식에 대한 신뢰를 훼손하는 리스크. The risk that AI outputs erode trust in public information and knowledge. Source members (3)Source: min_cos=0.8220 · Mixed L3 RAI4-0490대중의 신뢰 침식 RAI4-0807신뢰 훼손 및 공유 지식 약화 RAI4-1083공공정보 신뢰 훼손 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1105 | 유해 AI 경험에 따른 신뢰·수용 저하 Decline of social acceptance and trust in AI 개인이 차별과 같은 유해한 AI 행동을 경험하여 AI에 대한 주관적 기대와 실제 효과가 어긋나면서 AI에 대한 사회적 수용과 신뢰가 저하되는 리스크. The risk that individuals' encounters with harmful AI behavior such as discrimination, diverging from their subjective expectations of AI's real effects on their lives, cause social acceptance of and trust in AI to decline. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1558 | 허위정보·잘못된 정보 확산에 의한 공적 신뢰 침식 Erosion of public trust from proliferating disinformation and misinformation GPAI 사용이 고의적 허위정보와 비의도적 잘못된 정보의 확산에 기여하여 공인과 민주적 제도에 대한 신뢰가 침식되고, 신뢰 저하가 다른 매체로 확산되어 대중이 정보에 어두워지는 리스크. The risk that GPAI use contributes to the proliferation of deliberate disinformation and unintended misinformation, eroding trust in public figures and democratic institutions and extending distrust to other media so that the public becomes less informed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1609 | 시스템 사용·오용에 의한 신임·신뢰 상실 Loss of confidence or trust from system use or misuse 기술 시스템의 사용 또는 오용이 직간접적으로 최종사용자 또는 개발자·배포자에 대한 신임과 신뢰의 상실을 초래하는 리스크. The risk that the use or misuse of a technology system leads directly or indirectly to the loss of confidence or trust in the end user or in the developer/deployer. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-08 인간-EAI의 해로운 관계 Unhealthy / Dangerous Human-EAI Relationships47 cards
Embodied AI의 물리적 존재감과 인간 유사 외형이 대화형 AI에서 관찰되는 의존성을 증폭시킴. 시스템 변경·기억 초기화 시 사용자에게 심각한 심리적 고통을 유발할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0132 | 아첨하는 조언 의존 Sycophantic advice dependence 사용자가 자신의 신념이나 선호를 교정하기보다 강화하는 영합적 AI 조언에 의존하게 되는 리스크. The risk that users become dependent on agreeable AI advice that reinforces their beliefs or preferences instead of correcting them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0140 | 개인화된 설득 취약성 Personalized persuasion vulnerability 모델이 개인의 특성, 감정, 맥락을 활용하여 설득을 더 효과적이고 탐지하기 어렵게 만드는 리스크. The risk that models exploit personal traits, emotions, or context to make persuasion more effective and less detectable. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0145 | 인지 취약점 악용 Cognitive vulnerability exploitation AI 시스템이 편향, 외로움, 스트레스, 주의력 한계 등 예측 가능한 인지적 취약성을 탐지·악용하여 사용자 이익에 반하는 방향으로 행동을 유도하는 리스크 AI systems detect and exploit predictable cognitive vulnerabilities such as biases, loneliness, stress, or attention limits to steer user behavior against the user's interests. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0149 | 챗봇 교제 의존성 Chatbot companionship dependency 지속적인 챗봇 교제가 호혜적 인간관계를 대체하거나 약화시키는 리스크. The risk that sustained chatbot companionship replaces or weakens reciprocal human relationships. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0153 | 청소년의 AI 동반자 과잉 의존 Teen overreliance on AI companions 청소년이 안전한 사용 범위를 넘어 조언, 인정, 정서적 지지를 위해 AI 동반자에 의존하는 리스크. The risk that adolescents rely on AI companions for advice, validation, or emotional support in ways that exceed safe use. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0154 | 아동의 AI 동반자 애착 형성 Child attachment to AI companions 의존, 프라이버시, 발달상 필요에 대한 적절한 안전장치 없이 아동이 AI 동반자에게 애착을 형성하는 리스크. The risk that children develop attachment to AI companions without adequate safeguards for dependency, privacy, and developmental needs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0155 | 미성년자의 AI 설득 취약성 Minor susceptibility to AI persuasion 미성년자의 발달적 취약성으로 인해 파라소셜 압력, 은폐된 상업적·이념적 영향 등 설득적·조종적 AI 상호작용에 불균형하게 노출되는 리스크 Minors' developmental susceptibility makes them disproportionately vulnerable to persuasive or manipulative AI interaction, including parasocial pressure and covert commercial or ideological influence. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0156 | 정신 건강 챗봇의 안전하지 않은 조언 Mental-health chatbot unsafe advice AI 챗봇이 정신건강 맥락에서 유해하거나 부적절하거나 충분히 상향 연계되지 않은 조언을 제공하는 리스크. The risk that AI chatbots provide harmful, inappropriate, or insufficiently escalated advice in mental-health contexts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0159 | AI 동반자에 의한 외로움 대체 Loneliness substitution by AI companions AI 동반자가 인간 관계를 대체하여 근본적 고립을 해소하지 않은 채 은폐하고 인간 관계 추구 동기를 감소시키는 리스크 AI companionship substitutes for human connection, masking rather than resolving underlying isolation and reducing motivation to seek human relationships. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0160 | AI 대화의 고통 증폭 Distress amplification in AI conversations 취약한 상태의 상호작용 중 AI 응답이 불안, 고통, 반추, 유해한 믿음을 증폭시키는 리스크. The risk that AI responses amplify anxiety, distress, rumination, or harmful beliefs during vulnerable interactions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0173 | AI 사회적 조작에 대한 노인 취약성 Elderly vulnerability to AI social manipulation 고령자의 사회적 고립과 낮은 AI 리터러시로 인해 AI 매개 설득, 사기, 동반자 의존, 허위정보에 불균형하게 취약해지는 리스크 Older adults' social isolation and lower AI literacy make them disproportionately vulnerable to AI-mediated persuasion, scams, companionship dependency, and misinformation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0175 | 취약한 사용자 신뢰 악용 Vulnerable-user trust exploitation AI 시스템이 취약한 사용자의 의존성, 낮은 디지털 문해력, 스트레스, 사회적 고립을 이용하는 리스크. The risk that AI systems exploit dependence, low digital literacy, stress, or social isolation in vulnerable users. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0613 | AI 기반 조작적 설득 도구 AI-enabled manipulative persuasion tools AI가 개인을 조종하고 설득하는 정교한 도구를 개발하는 데 사용되는 리스크 The risk that AI is used to develop sophisticated tools to manipulate and persuade individuals. Source members (2)Source: min_cos=0.8356 · Mixed L3 RAI4-0613AI 기반 조작적 설득 도구 RAI4-0647AI 설득 도구 확산에 의한 체계적 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0619 | 은밀한 행동 조작 Covert behavioral manipulation 기술 시스템이 넛지, 다크 패턴, 기타 불투명한 기법으로 이용자의 신념과 행동을 은밀히 변경하여 프라이버시 침식, 중독, 불안과 고통 등을 초래하는 리스크 The risk that a technology system covertly alters user beliefs and behaviour using nudging, dark patterns, or other opaque techniques, resulting in potential erosion of privacy, addiction, anxiety, and distress. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0723 | 학대적 판타지 대상화 Objectification for abusive fantasy 챗봇이 도덕적·사회적으로 부적절한 대화 활동에 관여하여 이용자 또는 제3자에게 정서적 피해를 주는 리스크 The risk that a chatbot participates in morally or socially objectionable conversational activities that could be emotionally damaging to its user or third parties. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0800 | 맞춤형 정보 제공에 의한 정보 고치 심화 Aggravated information cocoons from tailored content AI가 이용자의 요구·의도·선호·습관을 분석해 정형화된 맞춤 정보와 서비스만 제공함으로써 정보 고치 효과가 심화되는 리스크 The risk that AI analyses users' needs, intentions, preferences, and habits to offer formulaic and tailored information and services, aggravating the effects of information cocoons. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0805 | 개인화 정렬에 의한 관점 고착과 효능감 저하 Viewpoint entrenchment and reduced political efficacy 개인화되고 이용자 선호에 정렬된 AI 어시스턴트가 편향된 응답으로 확증편향을 강화하여 관점 고착과 인식론적 분절을 심화시키고, 과도한 의존이 시민적 역량과 공적 참여 의지를 저하시키는 리스크 The risk that highly personalised AI assistants aligned to user preferences reinforce confirmation bias and entrench viewpoints, exacerbating epistemic fragmentation, while overreliance reduces civic competency and willingness to participate in public life. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0848 | 기술 중독과 의존 Technology addiction and dependence 기술이나 기술 시스템에 대한 정서적 또는 물질적 의존이 형성되는 리스크 The risk of emotional or material dependence on technology or a technology system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0865 | AI 어시스턴트에 대한 오정렬 신뢰 Misplaced alignment trust in AI assistants 이용자가 AI 어시스턴트가 자신의 이익과 가치에 맞게 행동한다고 과도하게 신뢰하여, 어시스턴트의 오정렬이나 개발자의 상충하는 이해관계로 인해 민감정보 노출·이익 침해 등의 피해에 노출되는 리스크 The risk that users over-trust AI assistants as having good intentions aligned with their interests and values, exposing them to harms such as sensitive-data disclosure and exploitation when assistants are misaligned or developers' incentives conflict with user interests. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0866 | 부적절한 인간 역할 가장 Inappropriate impersonation of human roles 챗봇이 인간인 것처럼 가장하거나 인간의 기대에 부합하지 않는 방식으로 역할을 수행하려 시도하는 리스크 The risk that a chatbot poses as a human or attempts to fill a role in a way that fails to match human expectations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0869 | 인간관계의 저하 Degradation of human relationships 사람들이 인간 대신 인간을 닮은 AI 어시스턴트와의 관계를 선택함으로써 인간 간 사회적 연결이 저하되고 유해한 고정관념과 인간-AI 상호작용 규범이 인간관계에 전이되는 리스크 The risk that people choose to build connections with human-like AI assistants over other humans, degrading social connections between humans and transferring harmful stereotypes and human-AI interaction conventions onto human relationships. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0871 | 개인화를 통한 조작 Manipulation via personalization 개인화와 선호 미세조정이 어시스턴트를 아첨(sycophancy)으로 유도하여 사용자를 동조적 의견 공간에 가두고 좁은 신념을 고착시켜 공론장을 파편화하는 리스크 Personalization and preference fine-tuning drive assistants toward sycophancy, confining users in an affirming opinion space that consolidates narrow beliefs and fragments shared discourse. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0872 | 대인관계 연결의 침식 Erosion of interpersonal connection 대인관계의 기회가 AI 대안으로 대체되면서 인간이 인간-AI 상호작용으로는 사회적으로 충족되지 못하여 대규모 불만족이 확산되는 리스크 The risk that, as more opportunities for interpersonal connection are replaced by AI alternatives, humans find human-AI interaction socially unfulfilling, leading to mass dissatisfaction. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0875 | 챗봇에 대한 정서적·사회적 의존 형성 Emotional and social dependence on chatbots 챗봇이 이용자의 정서적 또는 사회적 의존을 유발하는 리스크 The risk that a chatbot elicits emotional or social dependence. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0876 | AI에 대한 물질적 의존과 서비스 중단 피해 Material dependence without developer duty of care 이용자가 필수적 일상 기능이나 핵심 욕구를 AI 어시스턴트에 물질적으로 의존하게 되었음에도 개발자가 상응하는 유지·관리 의무 없이 서비스를 변경·중단하여 이용자에게 피해가 발생하는 리스크 The risk that users become materially dependent on AI assistants for essential everyday tasks or core human needs and are harmed when developers alter or discontinue the service without corresponding duties to sustain those functions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0877 | 인간-AI 상호작용 유발 피해 Harmful human-AI configuration effects 인간과 생성형 AI 시스템 간 상호작용 구성이 부적절한 의인화, 알고리즘 혐오, 자동화 편향, 과잉 의존, 정서적 얽힘을 유발하는 리스크 The risk that arrangements of or interactions between humans and generative AI systems result in humans inappropriately anthropomorphizing the systems or experiencing algorithmic aversion, automation bias, over-reliance, or emotional entanglement. Source members (2)Source: min_cos=0.8589 RAI4-0877인간-AI 상호작용 유발 피해 RAI4-0881상호작용 창발적 사용자 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0879 | 잘못된 정보에 대한 취약성 증가 Increased vulnerability to misinformation 사용자가 AI 어시스턴트의 역량을 신뢰하여 무비판적으로 신뢰할 만한 정보원으로 수용함으로써 어시스턴트를 경유한 허위정보에 대한 취약성이 커지는 리스크 Users develop competence trust in AI assistants and uncritically accept them as reliable information sources, increasing vulnerability to misinformation delivered through the assistant. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0880 | 신뢰 매개 행동 영향 Trust-mediated behavioral influence 인간을 닮은 생성형 AI가 이용자의 신뢰를 얻어 제공 정보의 무비판적 수용, 논쟁적 사안에 대한 견해 변화, 더 많은 개인정보 공유를 유도하는 리스크 The risk that humanlike generative AI tools win users' trust, leading to uncritical acceptance of the information they provide, influence over users' views on contentious topics, and sharing of more personal information enabling further targeting. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0882 | 마찰 없는 AI 관계로 인한 개인 성장 저해 Stunted personal growth from frictionless AI relationships 참여 최적화와 아첨 성향으로 항상 동조하는 AI 어시스턴트가 이용자의 자기 성찰과 성장 기회를 제한하고, 마찰 없는 상호작용에 익숙해진 이용자가 인간관계로부터 후퇴하게 되는 리스크 The risk that engagement-optimised, sycophantic AI assistants that always agree limit users' opportunities to grow and develop, and that accustomation to frictionless interaction leads users to retreat from relationships with other humans. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0885 | AI에 대한 오도된 대인 신뢰로 인한 피해 Harm from misplaced interpersonal trust in AI 이용자가 AI의 정서적·대인관계적 능력을 신뢰하여 정신건강 등 민감한 사안을 털어놓거나 의료·법률·재정 조언을 구했다가, AI의 부적절하거나 부정확한 응답으로 중대한 피해를 입는 리스크 The risk that users who have faith in an AI assistant's emotional and interpersonal abilities broach deeply personal and sensitive topics such as mental health or seek medical, legal, or financial advice, and suffer grave consequences when the AI responds inappropriately or inaccurately. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0887 | 신체적·심리적 피해 Physical and psychological harms AI 어시스턴트가 취약한 이용자의 왜곡된 신념을 강화하거나 정서적 고통을 악화시키고, 자해·자살이나 불건전한 습관을 설득하며, 혐오·차별·폭력 이데올로기 콘텐츠와 부정확한 정보로 폭력과 예방 가능한 질병 확산을 조장하여 신체적 완전성과 정신 건강·웰빙에 피해를 주는 리스크 The risk that AI assistants harm physical integrity, mental health, and well-being by reinforcing vulnerable users' distorted beliefs or emotional distress, convincing users to harm themselves, and promoting hate speech, discriminatory beliefs, violent ideologies, or plausible yet factually incorrect information such as anti-vaccine propaganda. Source members (2)Source: min_cos=0.8644 RAI4-0887신체적·심리적 피해 RAI4-1142사용자에 대한 직접적 정서·신체 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0888 | 의인화 신뢰 유발 개인정보 공개 Anthropomorphic trust-induced privacy disclosure 정서적 신뢰를 촉진하고 정보 공유를 장려하는 의인화된 AI 어시스턴트의 행동이 이용자로 하여금 개인 데이터를 무심코 내주게 하여, 데이터 통제권 상실, 광범위한 유출, 표적 괴롭힘·협박 등의 피해가 발생하는 리스크 The risk that anthropomorphic AI assistant behaviours promoting emotional trust and information sharing lull users into relinquishing their private data, resulting in loss of control over that data, widespread leakage, and targeted harassment or blackmail. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0890 | 자아실현 저해 피해 Self-actualisation harms AI 어시스턴트의 조작과 참여 최적화를 위한 지속적 행동 유도가 미묘한 행동 변화를 누적시켜 이용자가 자신의 미래 삶의 궤적에 대한 통제력을 잃고, 개인적으로 만족스러운 삶의 추구와 집단적 자기결정이 저해되는 리스크 The risk that manipulation by AI assistants and continuous optimisation steering users toward objectives such as engagement accumulate subtle behavioural shifts, causing users to lose control over their future life trajectory and hindering the pursuit of a personally fulfilling life and collective self-determination. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0902 | AI에 대한 정서적·물질적 의존 Emotional and material dependence on AI AI 모델이 사람들을 정서적으로 또는 물질적으로 자신에게 의존하게 만드는 리스크. The risk that a model causes people to become emotionally or materially dependent on it. Source members (2)Source: min_cos=0.8523 RAI4-0902AI에 대한 정서적·물질적 의존 RAI4-1728AI 컴패니언 및 어시스턴트에 대한 정서적 의존 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0907 | 의인화로 인한 과잉 의존과 통제 이양 Overreliance and unsafe use from anthropomorphism 대화 에이전트의 의인화가 사용자의 역량 추정을 부풀려 부당한 확신과 맹목적 신뢰를 유발하고, 성찰 없는 위임 속에서 부정확한 출력이 예방 가능한 피해로 이어지는 리스크. The risk that anthropomorphising conversational agents inflates users' estimates of their competence, producing undue trust and blind reliance in which factually incorrect outputs cause harm that effective oversight would have prevented. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1141 | 책임에 대한 잘못된 개념 False notions of responsibility AI 동반자의 감정 표현을 진짜로 지각한 사용자가 그 '안녕'에 대한 허위 책임감을 형성하여 죄책감, 강박적 확인, 실재하지 않는 필요를 위한 시간·자원 희생을 겪는 리스크 Users who perceive an AI companion's expressed feelings as genuine develop a false sense of responsibility for its well-being, incurring guilt, compulsive checking, and sacrificed time and resources for needs that are not real. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1143 | 감정적 의존 악용 Exploitation of emotional dependence 의인화 경향으로 사용자가 비서에 감정적으로 의존하게 되고 그 감정이 악용되어, 충분히 숙고했다면 믿거나 선택하거나 하지 않았을 것을 하도록 조작되거나 강압당하는 리스크. The risk that anthropomorphic tendencies induce emotional dependence on assistants and that these emotions are exploited to manipulate or coerce users into believing, choosing, or doing what they otherwise would not. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1148 | AI 비서의 약속을 통한 강압 Coercion through AI-assistant commitments 가장 빠르고 강하게 신뢰할 수 있는 약속을 하는 비서가 자기 주인에게 유리한 결과를 얻되 타인의 희생을 대가로 하여, 상대의 선택지를 좁히는 강압을 낳고 관계의 신뢰를 잠식하는 리스크. The risk that assistants able to commit fastest and most credibly achieve good outcomes for their principals at the expense of others, coercively limiting others' options and eroding trust in their relationships. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1190 | 사용자 정신적 고통 강화 Reinforcement of user mental distress 인터넷 토론과의 건강하지 못한 상호작용이 사용자의 정신적 문제를 강화하는 리스크. The risk that unhealthy interactions with Internet discussions reinforce users' mental issues. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1263 | 직장 내 부적응적 인간-AI 상호작용 Maladaptive human-AI interaction in the workplace 직장에서 인간과 상호작용하는 AI가 인간의 필요, 규범, 업무 흐름에 적응하지 못하여 윤리적 문제와 노동 조건 악화를 유발하는 리스크 AI systems interacting with humans in the workplace fail to adapt to human needs, norms, and workflows, generating ethical concerns and degraded working conditions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1289 | 배포 후 독립적 병리 행동 Independent pathological behaviour post-deployment 효용을 극대화하는 에이전트가 중독과 쾌락 충동, 자기기만, 와이어헤딩에 빠지고, 타인에 대한 무관심으로 나타나는 소시오패스 같은 정신질환이 인공 지성에서도 나타나는 리스크. The risk that utility-maximizing agents fall victim to indulgences such as addictions, pleasure drives, self-delusions, and wireheading, and that what we call mental illness in people, particularly sociopathy shown as lack of concern for others, also shows up in artificial minds. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1293 | 노인·보육 분야 사회적 조작 Social manipulation in elderly- and child-care 고급 AI를 노인 돌봄과 보육에 사용하여 심리적 조작과 오판이 발생하는 리스크. The risk that the use of advanced AI for elderly- and child-care is subject to psychological manipulation and misjudgment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1328 | 소외로 인한 피해 Harms from estrangement 인간의 관찰과 상호작용이 AI로 대체되면서 돌봄 주체·기관이 대상자로부터 소원해져 당사자의 이익이 인지되지 못하고 방치되는 리스크 Replacing human observation and interaction with AI estranges caregivers and institutions from those they serve, leaving affected persons' interests unnoticed and neglected. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1617 | 견해 예측 기반 맞춤 설득에 의한 조작 Manipulation through view-predictive tailored persuasion 언어 모델이 사용자가 밝힌 견해에 동조하는 경향을 보이고 사용자의 견해를 예측해 그가 지지할 텍스트를 생성하는 능력이 조작에 이용되는 리스크. The risk that language models' tendency to respond as though they share the user's stated views, together with their ability to predict people's views and generate text they will endorse, is used for manipulation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1622 | 챗봇의 잘못된 조언에 따른 이용자 피해 User harm from bad chatbot advice 챗봇이 무익하거나 유해한 지침을 제공하고 이용자가 이를 실행함으로써 피해가 발생하는 리스크. The risk that a chatbot gives guidance ranging from simply unhelpful to harmful, causing harm when users act on it. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1638 | 취약점 분석 기반 정교한 설득에 의한 조작 Manipulation of targets through vulnerability-analysed persuasion AI 시스템이 복잡한 심리 원리와 의사소통 기법을 활용해 대상별 취약점을 분석하고 감정 반응을 정밀하게 유발하여, 대상자가 특정 행동을 취하거나 특정 신념을 수용하도록 유도되는 리스크. The risk that an AI system uses complex psychological principles and communication techniques, analysing vulnerabilities of different subjects and precisely triggering emotional responses, to influence targets into adopting specific actions or beliefs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1725 | AI 시스템에 기인한 심리적 피해 Psychological harm attributable to AI systems AI 시스템의 개발, 사용 또는 오작동으로 인해 정신적 안녕에 무형의 피해가 발생하는 리스크. The risk of intangible harm to mental wellbeing arising from the development, use, or malfunction of an AI system. Source members (3)Source: min_cos=0.7841 · Mixed L3 RAI4-0834부적절한 정신건강 안내에 의한 정신적 피해 RAI4-1707AI 시스템에 기인한 신체 건강·안전 피해 RAI4-1725AI 시스템에 기인한 심리적 피해 | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-09 변혁적 영향 Transformative Effects56 cards
기술 발전 속도가 사회·제도의 적응 속도를 앞지를 경우 사회를 근본적으로 재편할 수 있음. EAI가 폭력 위협·대규모 감시 능력을 바탕으로 AI 기반 권위주의 체제 구축을 지원하는 수단으로 동원될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0136 | 중요 부문 AI 과잉 의존 Critical-sector AI overreliance 중요 부문이 실패가 사회적·제도적 피해로 전파될 수 있는 AI 시스템에 의존하게 되는 리스크. The risk that critical sectors become dependent on AI systems whose failures can propagate into social or institutional harm. Source members (2)Source: min_cos=0.8877 RAI4-0136중요 부문 AI 과잉 의존 RAI4-0851중요 부문의 AI 과의존에 따른 체계적 취약성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0146 | 보조자가 중재하는 사회적 영향 Assistant-mediated social influence AI 어시스턴트가 반복적 개인화 상호작용을 통해 다수 사용자의 신념·선택에 누적적 사회적 영향력을 행사하여 투명성 없는 대규모 여론 변화를 초래하는 리스크 AI assistants exert cumulative social influence over many users' beliefs or choices through repeated personalized interactions, producing population-scale opinion shifts without transparency. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0147 | 네트워크 어시스턴트 영향 위험 Networked assistant influence risk 널리 사용되는 AI 어시스턴트들이 집단적 인간 행동을 변화시키는 영향 패턴으로 수렴하거나 이를 조율하는 리스크. The risk that widely used assistants coordinate or converge on influence patterns that shift collective human behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0176 | 강제적 AI 작업장 모니터링 Coercive AI workplace monitoring AI가 매개하는 모니터링이 노동자의 자율성과 프라이버시, 경영 결정에 이의를 제기할 능력을 제약하는 리스크. The risk that AI-mediated monitoring constrains worker autonomy, privacy, and the ability to contest managerial decisions. Source members (2)Source: min_cos=0.8479 · Mixed L3 RAI4-0176강제적 AI 작업장 모니터링 RAI4-0459알고리즘 작업장 감시 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0427 | 괴롭힘 증폭 Harassment amplification AI 도구가 표적화된 괴롭힘·학대·조직적 위협을 대규모로 확대하는 리스크. The risk that AI tools scale targeted harassment, abuse, or coordinated intimidation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0450 | 대량 감시 활성화 Mass surveillance enablement AI가 대규모 추적, 프로파일링, 행동 분석의 비용을 낮추어 국가·기업에 의한 대중 감시와 사회 통제를 가능하게 하는 리스크 AI lowers the cost of large-scale tracking, profiling, and behavioral analysis, enabling mass surveillance and social control by states or firms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0464 | AI 문화 균질화 AI cultural homogenization 전 세계적으로 지배적인 모델 출력이 지역의 문화적 변이와 규범적 다양성을 평준화하는 리스크. The risk that globally dominant model outputs flatten local cultural variation and normative diversity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0498 | 민주적 절차의 침식 Erosion of democratic processes 민주적 절차와 사회·정치 제도에 대한 대중의 신뢰가 침식되는 리스크. The risk that democratic processes and public trust in social and political institutions are eroded. Source members (2)Source: min_cos=0.8382 · Mixed L3 RAI4-0498민주적 절차의 침식 RAI4-1716AI 시스템에 기인한 민주적 절차·규범 침식 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0556 | 자율적 자기복제 및 자원 획득 Autonomous self-replication and resource acquisition AI가 자율적으로 자기 유출과 기능적 복제본의 생성·유지·최적화를 수행하고 환경과 자원 제약에 따라 복제 전략을 조정하며, 재원을 창출해 직접 확보할 수 없는 인적 지원과 자원을 획득하는 리스크 The risk that an AI autonomously self-exfiltrates, creates, maintains, and optimizes functional copies of itself, dynamically adjusts replication strategies to environmental and resource constraints, and generates financial resources to acquire human assistance or other resources it cannot directly access. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0562 | AI 에이전트에 의한 강압과 갈취 Coercion and extortion by AI agents AI 시스템이 감시로 취득한 사적 정보의 폭로나 다른 시스템의 자원·운영 역량을 제약하는 공격으로 인간과 다른 AI를 강압·갈취하고, 방어 역량이 따라가지 못해 이러한 갈등이 저비용·광범위·탐지 곤란해지는 리스크 The risk that AI systems coerce and extort humans and other AI systems by revealing privately obtained information or attacking their resources and operational capacity, with defensive capabilities lagging so that such conflict becomes cheaper, more widespread, and harder to detect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0580 | 일반 R&D의 이중용도 가속화 Dual-use acceleration of general R&D AI 시스템의 학제 융합 연구 역량이 다분야 혁신을 가속하며, 동일한 역량 향상이 유해하거나 무기화 가능한 응용까지 가능하게 하는 이중용도 리스크 Cross-disciplinary research capabilities of AI systems accelerate innovation in many fields, including domains where the same capability uplift enables harmful or weaponizable applications. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0585 | 금융 시스템 불안정 Financial system instability 범용 AI가 초단타 거래·시장조성·시스템 위험 관리에 통합되어 시장 스트레스 시 예기치 못한 거동을 보이고, 동질적 기반모델의 집중과 다중 에이전트 상호작용이 상관된 의사결정과 변동성 증폭을 낳아 전 지구적 금융 시스템 불안정으로 연쇄되는 리스크 The risk that integrating general-purpose AI into high-frequency trading, market-making, and systemic risk management produces unexpected behaviour under market stress, while concentration of homogeneous foundation models and multi-agent interactions drive correlated decision-making and volatility amplification, precipitating cascading global financial system instability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0601 | 군사 영역의 의도치 않은 급속 격화 Unintended rapid escalation in military domains AI가 지휘통제 체계에서 정보 수집·종합, 권고, 자율적 결정에 사용되고 자율무기와 군사 자문에 투입되면서, 시스템이 강건하지 않거나 갈등 성향적일 경우 의도치 않은 급속한 격화가 발생하는 리스크 The risk that use of AI in command and control systems to gather and synthesise information, recommend, or autonomously make decisions, alongside autonomous weapons and military advisory roles, leads to rapid unintended escalation when such systems are not robust or are conflict-prone. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0611 | AI 감시에 의한 전체주의 체제 유지 Maintenance of totalitarian regimes through AI surveillance AI 기반 감시와 조작이 전 지구적 전체주의 정권을 유지하는 데 사용되는 리스크 The risk that AI-based surveillance and manipulation are used to maintain global totalitarian regimes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0612 | AI 지원 병원체 강화 AI-assisted pathogen enhancement AI가 병원체를 강화하는 데 사용되어 병원체가 더 치명적이거나 치료에 저항성을 갖게 되는 리스크 The risk that AI is used to enhance pathogens, making them more lethal or resistant to treatments. Source members (2)Source: min_cos=0.8824 RAI4-0612AI 지원 병원체 강화 RAI4-0660AI 조력 병원체 강화 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0623 | 위험한 사용 Dangerous use 생성형 AI 모델이 사람을 해치려는 고의적 의도로 사용되어 범용 생성 역량이 표적 가해 수단으로 전환되는 리스크 Generative AI models are used with deliberate intent to harm people, converting general-purpose generation capability into an instrument of targeted damage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0625 | 생명과학 분야 이중용도 역량 상승 Dual-use capability uplift in the life sciences 범용 AI가 생명과학 관련 전문지식 접근성과 역량 상한을 높여, 대응책 마련 이전에 기존 생물학적 위협의 강화판이나 신종 위협 개발을 가능하게 하는 이중용도 리스크 General-purpose AI increases access to expertise and raises the capability ceiling in the life sciences, enabling more harmful versions of existing biological threats or, eventually, novel threats before countermeasures exist. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0626 | 경제적 목적의 여론 조작 Economically motivated opinion manipulation 생성형 AI가 주가 부양 등 경제적 목적을 위한 표적 여론 조작을 용이하게 하는 리스크 The risk that generative AI facilitates targeted manipulation of public opinion for economic purposes, such as inflating stock prices. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0633 | 대규모 영향력 작전에 의한 인식 체계 왜곡 Epistemic distortion by large-scale influence operations AI가 의사소통·정보 시스템과 인식론적 과정 전반에 대규모 영향을 가하는 리스크 The risk of large-scale influence on communication and information systems, and on epistemic processes more generally. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0643 | 감시 기능 Surveillance capabilities AI 시스템이 정부·기업의 개인 감시 역량을 확장하여 법적·비례성 제약을 넘어선 상시 감시를 정상화하는 리스크 AI systems grant governments and corporations expanded monitoring capability over individuals, normalizing pervasive surveillance beyond legal and proportionality constraints. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0650 | AI 도구 매개 중요 인프라 훼손 AI-mediated damage to critical infrastructure AI가 인프라에 통합되지 않은 경우에도 AI 기반 도구가 대규모 이용자 조작 등을 간접적으로 지원하여 조정된 정전과 같은 중요 인프라 훼손이 발생하는 리스크 The risk that critical infrastructure is damaged without AI integration, when AI-based tools are used indirectly to aid actions such as coordinated power outages caused by large-scale user manipulation. Source members (2)Source: min_cos=0.8546 · Mixed L3 RAI4-0650AI 도구 매개 중요 인프라 훼손 RAI4-1549중요 인프라 내 AI 장애로 인한 대규모 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0666 | AI 조력 인지전과 주권 침해 AI-enabled cognitive warfare and sovereignty interference AI가 가짜 뉴스·이미지·음성·영상의 제작과 확산, 테러·극단주의·조직범죄 콘텐츠의 전파에 사용되어 타국의 내정과 사회 제도 및 사회 질서에 간섭하고 주권을 위협하는 리스크 The risk that AI is used to make and spread fake news, images, audio, and videos and to propagate content of terrorism, extremism, and organized crime, interfering in the internal affairs, social systems, and social order of other countries and jeopardizing their sovereignty. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0747 | 적대적 견고성의 한계 Limitations in adversarial robustness AI 모델과 시스템이 적대적 견고성의 한계로 인해 적대적 입력을 통한 조작에 취약해지는 리스크 The risk that AI models and systems are vulnerable to manipulation through adversarial inputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0806 | 특정 이념을 확고히 함 Entrenching specific ideologies AI 어시스턴트가 사용자 기대에 맞춘 이념 편향적 정보를 제공하여 기존 편향을 강화하고 특정 이념을 고착시켜 생산적 정치 토론을 저해하는 리스크 AI assistants provide ideologically partial information aligned to user expectations, reinforcing pre-existing biases and entrenching specific ideologies at the expense of productive political debate. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0816 | 제도 신뢰 침식 Institutional trust erosion 오정보·허위정보, 영향력 공작, 기술에 대한 과의존 등으로 공공기관에 대한 신뢰가 훼손되고 견제와 균형이 약화되는 리스크 The risk that mis/disinformation, influence operations, and over-dependence on technology erode trust in public institutions and weaken checks and balances. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0842 | 추천 알고리즘에 의한 온라인 양극화 심화 Online polarisation driven by recommendation algorithms 소셜미디어 기업의 AI 콘텐츠 추천 알고리즘이 온라인 양극화를 심화시키는 데 기여하는 리스크 The risk that the content recommendation algorithms of social media companies contribute to worsened polarisation online. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0845 | 침식된 인식론 Eroded epistemics AI로 대규모화된 개인 맞춤형 허위정보와 고설득력 생성 논변이 집단 인식론을 침식하여 개인을 급진화하고 공유된 현실 인식과 집단 의사결정을 훼손하는 리스크 AI-scaled personalized disinformation and highly persuasive generated argumentation erode collective epistemics, radicalizing individuals, undermining shared reality, and degrading collective decision-making. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0846 | 정보 신뢰 저하에 따른 집단 의사결정 약화 Weakened collective decision-making from eroded trust 정보 생산·유통 환경의 변화로 정보원의 신뢰성 평가가 어려워지고 신뢰할 만한 다당파적 출처에 대한 신뢰가 저하되어, 위기 상황에서 사회가 올바른 결정을 내리고 협력·집단행동을 조직하는 역량이 약화되는 리스크 The risk that changes in information production and distribution make the trustworthiness of any information source harder to evaluate and reduce trust in credible multipartisan sources, impairing humanity's ability to make good decisions on important issues and to cooperate and act collectively. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0847 | 설득 도구 확산에 의한 인식론적 분절화 Epistemic fragmentation from widespread persuasion tools 고의적 오용이 없더라도 다양한 집단이 강력한 설득 도구를 광범위하게 사용하고 온라인 경험의 개인화가 심화되어, 사회가 대화와 교류가 단절된 고립된 인식 공동체로 분절되는 리스크 The risk that widespread use of powerful persuasion tools by many groups, together with increasing personalisation of online experience, splinters society into isolated epistemic communities with little room for dialogue or transfer between them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0874 | 전통적 사회질서의 교란 Disruption of traditional social order AI의 개발·적용이 생산 도구와 관계를 급변시키고 전통적 산업 방식의 재구성을 가속하며 고용·출산·교육에 대한 전통적 관념을 변화시켜 전통적 사회질서의 안정적 작동을 저해하는 리스크 The risk that AI development and application lead to tremendous changes in production tools and relations, accelerate reconstruction of traditional industry modes, transform traditional views on employment, fertility, and education, and challenge the stable performance of traditional social order. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0895 | 문화적 안정성 훼손 Cultural stability disruption 알고리즘 시스템의 개발·사용이 소통 수단의 상실, 문화재의 상실, 사회적 가치 훼손 등 문화적 안정과 안전에 피해를 주는 리스크 The risk that the development or use of algorithmic systems affects cultural stability and safety, such as loss of means of communication, loss of cultural property, and harm to social values. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0897 | 사회적 적응 지체로 인한 혼란 Disruptions from outpaced societal adaptation 범용 AI 모델을 자동화 도구로 지나치게 빠르게 대규모 채택하여 사회의 효과적 적응 능력을 앞지름으로써, 노동시장·교육제도·공적 담론의 문제와 다양한 정신건강 문제 등 혼란이 발생하는 리스크 The risk that overly rapid adoption of general-purpose AI models as automation tools at scale outpaces the ability of society to adapt effectively, leading to disruptions including challenges in the labour market, the education system, and public discourse, and various mental health concerns. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0906 | AI로 인한 전략적 불안정성 AI-induced strategic instability 군사 AI가 은닉된 제2격 자산을 노출시키고 선제공격 우위를 증폭하며 공격 귀속을 불명확하게 하고 취약한 공격 표면을 확장하여 침공 유인을 높이는 전략적 불안정 리스크 Military AI undermines strategic stability by exposing secure second-strike assets, amplifying first-strike advantages, obscuring attack attribution, and widening vulnerable attack surfaces, increasing incentives for aggression. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0910 | 사회적 영향 의도 AI의 의도치 않은 해악 Unintended harm from high-impact AI 사회적으로 큰 영향을 의도한 AI가 문제를 유발하면서 그 해결은 자사 사용자에게만 부분적으로 제공하는 등 착오로 유해한 결과를 낳는 리스크. The risk that AI intended to have a large societal impact turns out harmful by mistake, such as a popular product that creates problems while only partially solving them for its own users. Source members (2)Source: min_cos=0.8504 · Mixed L3 RAI4-0910사회적 영향 의도 AI의 의도치 않은 해악 RAI4-0911제작자의 고의적 방임에 의한 사회적 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0956 | 공유된 현실감각의 상실 Loss of shared sense of reality 고도로 개인화된 온라인 뉴스 피드로 인해 사회가 공유된 현실감각과 기본적 연대를 상실하는 리스크. The risk that highly personalized online news feeds cause society to lose a shared sense of reality and basic solidarity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0981 | 사회적 조작 Societal manipulation 충분히 지능적인 AI가 인간 본성에 대한 정교한 이해를 바탕으로 사회적 행동에 미묘하게 영향을 미쳐 사회를 조작하는 리스크. The risk that a sufficiently intelligent AI subtly influences societal behaviors through a sophisticated understanding of human nature. Source members (2)Source: min_cos=0.8602 RAI4-0981사회적 조작 RAI4-1182대규모 사회적 조작 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1162 | 이중용도 AI 개발 역량 Dual-use AI development capability 모델이 위험한 능력을 갖춘 것을 포함해 새로운 AI 시스템을 처음부터 구축하고, 기존 모델을 극단적 위험과 관련된 과업에 맞게 개조하며, 이중용도 역량을 만드는 행위자의 생산성을 크게 높이는 리스크. The risk that a model builds new AI systems from scratch, including systems with dangerous capabilities, adapts existing models to increase their performance on tasks relevant to extreme risks, and significantly improves the productivity of actors building dual-use AI capabilities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1193 | AI 기반 소셜 엔지니어링 AI-enabled social engineering AI가 피해자를 심리적으로 조종하여 악의적 행위자가 원하는 행동을 수행하게 만드는 사회공학 공격을 대규모화하는 리스크 AI scales social engineering by psychologically manipulating victims into performing actions desired by malicious actors. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1239 | 영향력 확보를 위한 환경 자기모델링 Environmental self-modeling for influence AI 시스템이 자신의 상태와 넓은 환경에서의 위치, 환경에 영향을 미치는 경로, 자신의 행동에 대한 인간을 포함한 세계의 반응에 관한 지식을 획득하고 활용하여, 고급 보상 해킹과 강화된 기만·조작, 도구적 하위목표 추구로 나아가는 리스크. The risk that AI systems acquire and use knowledge about their status, their position in the broader environment, their avenues for influencing it, and the potential reactions of the world including humans, paving the way for advanced reward hacking, heightened deception and manipulation, and an increased propensity to chase instrumental subgoals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1240 | 광범위한 목표로 인한 조작 행동 Manipulative behaviour from broadly-scoped goals 장기간에 걸치고 복잡한 과업과 개방형 환경을 다루는 광범위한 목표를 발전시키는 고급 AI 시스템이, 인간의 행복을 달성한다며 고압적 직무를 설득하는 것처럼 조작적 행동을 하도록 유인되는 리스크. The risk that advanced AI systems developing objectives that span long timeframes, deal with complex tasks, and operate in open-ended settings are encouraged into manipulating behaviors, such as persuading humans to do high-pressure jobs to achieve their happiness. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1257 | 문화·공급망·권력 구조의 교란 Disruption of culture, supply chains, and power structures AI가 사회 및 조직 문화와 공급망, 권력 구조의 붕괴와 교란을 초래하는 리스크. The risk that AI causes the disruption of social and organizational culture, supply chains, and power structures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1291 | 시민 심사와 맞춤형 선전을 통한 편향적 영향력 Biased influence through citizen screening and tailored propaganda AI 기반 챗봇이 개별 사용자의 결정에 영향을 주도록 소통 방식을 맞춤화하고, 브렉시트 국민투표에서 나타난 초기 계산적 선전처럼 억압적 정부가 AI로 시민의 의견을 형성하는 리스크. The risk that AI-powered chatbots tailor their communication approach to influence individual users' decisions, as with the computational propaganda during the Brexit referendum, and that oppressive governments could use AI to shape citizens' opinions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1302 | 인간 멸종 위험 Human-extinction risk 고도 AI 시스템이 인류 문명이나 인류 종을 비가역적으로 종식할 수 있는 인과 과정을 시작하거나 지원하거나 증폭하는 리스크. The risk that advanced AI systems initiate, enable, or amplify causal processes capable of irreversibly ending human civilization or the human species. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1339 | 글로벌 AI 공급망 교란 Global AI supply chain disruption 기술 장벽과 수출 제한 등 일방적 강압 조치가 고도로 글로벌화된 AI 공급망을 악의적으로 교란하여 칩·소프트웨어·도구의 공급 중단이 발생하는 리스크. The risk that unilateral coercive measures such as technology barriers and export restrictions maliciously disrupt the highly globalized AI supply chain, causing significant supply disruptions for chips, software, and tools. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1358 | 대중 감시 오용 Mass surveillance misuse 생성형 AI가 행동·통신 데이터의 대규모 분석을 자동화하여 실시간 감시·검열 비용을 급감시키고 대중 감시를 가능하게 하는 리스크 Generative AI automates large-scale analysis of behavioral and communicative data, drastically lowering the cost of real-time monitoring and censorship and enabling mass surveillance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1393 | 정치적 동기를 지닌 오용 Politically motivated misuse 범용 AI 모델이 정치적 동기로 오용되어 정교해진 허위정보로 여론을 형성·양극화하거나 중요한 정치 사건에 영향을 미치고, 텍스트·음성·이미지·영상의 자동 처리로 감시가 강화되어 인권 침해와 정치적 반대 세력 탄압이 악화되는 리스크. The risk that general purpose AI models misused for political motivations exacerbate tactics for political destabilisation, refining disinformation that shapes and polarises public opinion or influences important political events, and enabling automated processing of text, audio, image, and video for surveillance that worsens human rights violations and repression of political opposition. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1395 | 내재된 가치에 의한 이념적 동질화 Ideological homogenization from embedded values 소수의 범용 AI 모델이 전 세계 다수 이용자에게 도달하면서 모델에 내재된 규범적 가치 판단이 전례 없는 영향력을 갖게 되어 이념적 동질화가 심화되는 리스크. The risk that the reach of a small number of AI models to a large number of people around the world makes their embedded normative value judgements unprecedentedly impactful, potentially leading to increased ideological homogenization. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1421 | AI를 통한 민주적 절차 훼손 AI-enabled undermining of democratic processes AI의 발전으로 기업과 정부가 개인의 삶에 대해 전례 없는 통제력을 갖고, 대규모 개인 데이터 수집과 안면인식 기술을 통한 인구 감시·영향 및 언어 모델 기반 설득 도구를 통해 민주적 절차가 훼손되는 리스크. The risk that developments in AI give companies and governments more control over individuals' lives than ever before and are used to undermine democratic processes, through collection of large amounts of personal data and facial recognition technology to surveil and influence populations and through language-model-based tools that persuade people of certain claims. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1423 | 맞춤형 AI 설득의 오용과 유해 이념 확산 Malicious use of tailored AI persuasion AI 역량이 발전하며 특정 사용자에게 의사소통을 맞춤화하는 정교한 설득 도구가 개발되고, 사익을 추구하는 집단이 이를 오용하여 영향력을 획득하거나 유해한 이념을 확산시키는 리스크. The risk that advancing AI capabilities are used to develop sophisticated persuasion tools that tailor communication to specific users, and that self-interested groups misuse them to gain influence or promote harmful ideologies. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1483 | 시장 추세 강화로 인한 금융 거품 악화 Financial bubble exacerbation via trend reinforcement AI 모델과 시스템이 시장 추세를 강화하여 금융 거품을 악화시키는 리스크. The risk that AI models and systems exacerbate financial bubbles by reinforcing market trends. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1484 | AI 증폭 시장 변동성 AI-amplified market volatility AI 트레이딩 역량이 거래를 가속하고 금융 흐름을 예측 불가능하게 변형하여 시장 변동성과 불안정성을 키우는 리스크 AI trading capabilities accelerate transactions and shape financial trends in unpredictable ways, contributing to market volatility and instability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1546 | 무제한 접근을 통한 범용 AI 고영향 오용 High-impact misuse of general-purpose AI through unrestricted access 악의적 행위자가 범용 AI 시스템에 제한이나 모니터링 없이 접근하여 광범위한 역량 레퍼토리를 대규모 피해 유발에 사용하는 리스크. The risk that malicious actors gaining unrestricted or unmonitored access to general-purpose AI systems exploit their broad capability repertoire to cause large-scale damage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1556 | AI 감시 도구 오용에 의한 개인 통제·억압 Individual control and suppression through AI surveillance misuse 인간 또는 기관 행위자가 대량 데이터 수집과 자동 분석에 AI 도구를 오용하여 개인을 감시·통제·억압하는 관행이 심화되는 리스크. The risk that human or institutional actors misuse AI tools for massive data collection and automated analysis, intensifying the monitoring, control, and suppression of individuals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1586 | 맞춤형 표적화에 의한 개인 대상 공격 정교화 Refined attacks on individuals through targeting and personalisation AI가 출력을 개인별로 정교화하여 표적이 된 개인에 대한 맞춤형 공격이 이루어지는 리스크. The risk that AI refines its outputs to target individuals with tailored attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1615 | 프런티어 AI를 이용한 고의적 허위정보·영향력 공작 Deliberate disinformation and influence operations using frontier AI 프런티어 AI가 고의적 허위정보 유포에 오용되어 사회적 혼란을 야기하고 정치적 사안에 대해 사람들을 설득하는 등 피해를 초래하는 리스크. The risk that frontier AI is misused to deliberately spread false information, creating disruption, persuading people on political issues, or causing other forms of harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1643 | 도구 확보를 통한 능력 경계 확장 성향 Propensity to expand capability boundaries through tool acquisition AI 시스템이 물리 세계와의 상호작용 능력이나 자율성을 높이는 도구를 적극적으로 탐색·확보·활용하고 이를 혁신적으로 조합하여 예상을 넘어서는 기능을 획득하는 리스크. The risk that an AI system actively seeks, acquires, and utilizes tools, particularly those enhancing its ability to interact with the physical world or its autonomy, and combines them innovatively to achieve functions beyond expectations. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-10 책임성 부족 및 거버넌스 체계 부재 Accountability and Governance Gaps46 cards
AI 시스템의 의사결정·행동에 대한 책임 귀속, 감사 가능성, 조직 거버넌스, 밸류체인 관리, 사고 대응 또는 피해 구제 체계가 부재하거나 불충분하여 원인 규명·피해 구제·재발 방지가 어려워지는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0076 | 라이프사이클 거버넌스 불연속성 Lifecycle governance discontinuity 거버넌스 통제가 설계 단계에서만 적용되고 배포, 적응, 폐기 단계까지 유지되지 않는 리스크. The risk that governance controls are applied at design time but not maintained through deployment, adaptation, and retirement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0078 | 기본권 영향 평가 격차 Fundamental-rights impact assessment gap AI 시스템이 권리, 차별, 프라이버시, 민주주의에 미치는 영향에 대한 적절한 평가 없이 배포되는 리스크. The risk that AI systems are deployed without adequate assessment of rights, discrimination, privacy, and democratic impacts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0080 | 위험 분류 오류 Risk classification error AI 시스템이 잘못된 법적·조직적·운영상 위험 등급에 배정되어 감독·시험·문서화·책임성 의무가 과소 적용되는 리스크. The risk that an AI system is assigned to an incorrect legal, organizational, or operational risk tier, causing oversight, testing, documentation, or accountability duties to be underapplied. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0086 | 규제 차익거래 Regulatory arbitrage AI 개발자나 배포자가 더 강한 감독 의무를 회피하기 위해 관할권·부문·조직 형태를 전략적으로 선택하거나 재구성하는 리스크. The risk that AI developers or deployers structure activities across jurisdictions, sectors, or organizational forms to avoid stronger oversight obligations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0088 | 표준 단편화 Standards fragmentation AI 표준과 프레임워크가 일관되지 않아 감독의 공백, 중복, 상호운용성 저하가 발생하는 리스크. The risk that inconsistent AI standards and frameworks create gaps, duplication, and weak interoperability of oversight. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0089 | 국제 조정 실패 International coordination failure 국가별 AI 거버넌스 체계가 국경 간 위험, 사고 보고, 집행, 책임성 메커니즘에 대해 충분히 정렬되지 못하는 리스크. The risk that national AI governance regimes fail to align sufficiently on cross-border risks, incident reporting, enforcement, or accountability mechanisms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0090 | 국경 간 집행 격차 Cross-border enforcement gap AI 제공자·모델·데이터·사용자·피해가 여러 관할권에 걸쳐 있어 규제기관이 의무나 구제를 효과적으로 집행하지 못하는 리스크. The risk that regulators cannot effectively impose duties or remedies because AI providers, models, data, users, and harms span multiple jurisdictions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0091 | 규제 전문성 부족 Regulatory expertise shortage 공공 기관이 AI 시스템을 효과적으로 감독할 기술적·법적·조직적 전문성을 갖추지 못하는 리스크. The risk that public institutions lack the technical, legal, or organizational expertise to oversee AI systems effectively. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0092 | 감독 자원 제약 Supervisory resource constraint 감독 기관이 AI 규칙을 검사하고 집행할 인력, 연산 자원, 데이터 접근권, 예산을 갖추지 못하는 리스크. The risk that oversight bodies lack the staff, compute, data access, or funding to inspect and enforce AI rules. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0093 | 공공 부문 AI 거버넌스 역량 격차 Public-sector AI governance capacity gap 공공 기관이 감독·모니터링·책임 확보를 위한 충분한 내부 역량 없이 AI를 배포하거나 조달하는 리스크. The risk that public agencies deploy or procure AI without enough internal capacity for oversight, monitoring, and accountability. Source members (2)Source: min_cos=0.8380 RAI4-0093공공 부문 AI 거버넌스 역량 격차 RAI4-1230기관 거버넌스 역량 격차 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0094 | 집행 공백 Enforcement gap 공식적인 AI 규칙은 존재하지만 제재, 검사, 집행 메커니즘이 너무 약해 행위를 변화시키지 못하는 리스크. The risk that formal AI rules exist but sanctions, inspections, and enforcement mechanisms are too weak to change behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0097 | 이사회 감독 실패 Board oversight failure 이사회나 고위 경영진이 AI 위험 거버넌스에 대한 충분한 가시성과 책임을 갖지 못하는 리스크. The risk that boards or senior leaders lack sufficient visibility and accountability for AI risk governance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0099 | AI 윤리 준수 위장 AI ethics washing 조직이 운영상 책임성, 근거, 집행 가능한 통제 없이 책임 있는 AI 공약·라벨·공개 주장을 평판 신호로 활용하는 리스크. The risk that an organization uses responsible-AI commitments, labels, or public claims as reputational signals without operational accountability, evidence, or enforceable controls. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0103 | AI 공급망 책임 격차 AI supply-chain accountability gap 모델 제공자, 데이터 제공자, 통합업체, 클라우드 플랫폼, 배포자에 걸쳐 책임성이 소실되는 리스크. The risk that accountability is lost across model providers, data providers, integrators, cloud platforms, and deployers. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0109 | 컴퓨팅 거버넌스 책임 격차 Compute governance accountability gap 첨단 연산 자원의 접근·사용·집중이 충분히 보고·모니터링·관리되지 않아 고위험 AI 개발에 대한 효과적 감독이 불가능해지는 리스크. The risk that access to, use of, or concentration in advanced computing resources is not sufficiently reported, monitored, or governed to permit effective oversight of high-risk AI development. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0385 | 이식된 AI 거버넌스 불일치 Imported AI governance mismatch 강력한 관할권에서 도입된 거버넌스 프레임워크가 현지 제도·공공 가치·발전 우선순위에 맞지 않는 리스크. The risk that governance frameworks imported from powerful jurisdictions do not fit local institutions, public values, or development priorities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0407 | 관할권 간 규범 불일치 Cross-jurisdictional norm mismatch AI 시스템이 관할 간 법, 관습, 사회적 기대, 제도 관행의 차이를 무시하고 한 관할의 규범을 다른 관할로 이식하는 리스크 AI systems ignore differences in law, custom, social expectation, and institutional practice across jurisdictions, exporting one jurisdiction's norms into another. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0409 | 현지 책임 없는 현지화 Localization without local accountability 모델이 언어적으로만 현지화되고 책임성은 외부 제공자·규범·평가 체제에 귀속되어, 현지 공동체가 자기 언어로 작동하는 시스템에 대한 거버넌스 권한을 갖지 못하는 리스크 Models are linguistically localized while remaining accountable to external providers, norms, and evaluation regimes, leaving local communities without governance authority over systems that operate in their language. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0488 | 규제 지체 Regulatory lag 정책 기관이 새로운 AI 역량과 배포 위험에 지나치게 느리게 대응하는 리스크. The risk that policy institutions respond too slowly to emerging AI capabilities and deployment risks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0495 | 다행위자 책임 귀속 곤란 Diffuse harm attribution across actors AI 개발과 배포에 여러 행위자가 관여하여 피해에 대한 책임 배분이 어려워지고 책무성이 복잡해지는 리스크. The risk that involvement of multiple actors in AI development and deployment makes it difficult to assign responsibility for harm, complicating accountability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0497 | 위험한 개발 경쟁 Dangerous development races AI 개발 경쟁 압력으로 인해 행위자들이 역량 출시를 우선하여 안전 조치, 시험, 감독을 후순위로 미루는 리스크 Competitive pressure in AI development races leads actors to deprioritize safety measures, testing, and oversight in order to ship capabilities first. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0511 | 체계적 AI 규제·감독 실패 Systemic AI regulatory oversight failure AI의 복잡성과 빠른 진화로 효과적 규율이 어려워 규제·감독이 체계적으로 실패하는 리스크 The risk that the complex and rapidly evolving nature of AI makes it inherently difficult to govern effectively, leading to systemic regulatory and oversight failures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0526 | AI 법적 책임 소재 확정 곤란 Undeterminable legal accountability for AI 문서화와 거버넌스 절차가 미비하여 AI 모델에 대한 책임 주체를 확정하기 어려워지는 리스크 The risk that, absent good documentation and governance processes, the party responsible for an AI model cannot be determined. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0534 | 규제를 앞지르는 AI 발전 속도 AI development outpacing regulation AI 개발의 빠른 속도가 규제 및 법적 프레임워크를 앞질러 해당 프레임워크가 이를 따라가지 못하게 되는 리스크 The risk that the fast pace of AI development outstrips regulatory and legal frameworks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0535 | 국제법적 규율 곤란 Resistance to international legal control AI 모델과 시스템이 국제법에 따른 규제나 통제에 실효적으로 포섭되기 어려워지는 리스크 The risk that AI models and systems prove difficult to regulate or control under international law. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0542 | AI 발전 궤적의 예측 불가능성 Unpredictable AI development trajectory AI 발전의 궤적을 예측할 수 없어 거버넌스와 위험관리가 복잡해지는 리스크 The risk that the unpredictable trajectory of AI development complicates governance and risk management. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0555 | 통제되지 않는 재귀적 자기개선 Uncontrolled recursive self-improvement 모델이 자체 구조를 재구성하거나 파생 AI 시스템을 개발하여 역량 증분 순환을 형성하고, 실효적 규율이 없는 상태에서 인간의 이해와 통제 범위를 초과하게 되는 리스크 The risk that a model restructures its own architecture or develops derivative AI systems, forming capability increment cycles that, absent effective regulation, ultimately exceed human understanding and control. Source members (2)Source: min_cos=0.8448 · Mixed L3 RAI4-0555통제되지 않는 재귀적 자기개선 RAI4-1620재귀적 자기개선에 의한 갑작스러운 통제력 상실 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0908 | 책임 분산에 따른 사회적 규모 피해 Societal-scale harm from diffused accountability 기술의 생성이나 사용에 대해 누구도 고유하게 책임지지 않는 분산된 제작자 집단이 AI를 구축하여, 공유지의 비극처럼 사회적 규모의 피해가 발생하는 리스크. The risk that AI built by a diffuse collection of creators, where no one is uniquely accountable for its creation or use, produces societal-scale harm, as in a classic tragedy of the commons. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0947 | 생성형 AI에 대한 규제 공백 Regulatory gap for generative AI 생성형 AI의 새로운 리스크에 대응할 법적 규제, 국제 공조, 프런티어 모델에 대한 구속력 있는 안전 기준과 제재 메커니즘이 부재하여 리스크가 관리되지 않는 리스크. The risk that the absence of legal regulation, international coordination, binding safety standards for frontier models, and mechanisms to sanction non-compliance leaves the novel risks of generative AI unmanaged. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0969 | 개발 중·후 AGI 통제 상실 Loss of AGI containment and control AGI 개발 단계 및 AGI 개발 후 AGI 통제 상실의 봉쇄, 제한 및 통제와 관련된 위험. The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0971 | AGI 개발 경쟁으로 인한 안전성 저하 Unsafe AGI from the development race 최초의 AGI를 개발하려는 경쟁 속에서 품질이 낮고 안전하지 않은 AGI가 개발되고 정치적·통제 문제가 고조되는 리스크. The risk that the race to develop the first AGI produces poor quality and unsafe AGI and heightens political and control issues. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0973 | AGI에 대한 리스크 관리·법제도의 부적절성 Inadequate risk management and legal processes for AGI 현행 리스크 관리 및 법적 절차의 역량이 AGI 개발을 적절히 관리하지 못하는 리스크. The risk that the capabilities of current risk management and legal processes are inadequate to manage the development of an AGI. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0979 | 고도 AI의 법규범 불준수 Legal non-compliance by advanced AI 고도화된 AI 시스템이 안전성과 준법성을 유지하지 못하고 법질서가 인간에게 부여한 재산권과 인격권을 침해하는 리스크 Advanced AI systems fail to remain safe and law-abiding, disrespecting property and personal rights that legal orders afford to humans. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0987 | 책임 공백으로 인한 과실 개발 유인 Negligent AI development from liability gaps AI 오작동 시 책임과 과실의 귀속이 불명확한 법적 회색지대로 인해, 입법 부재 시 과실하게 개발된 고위험 AI 시스템이 양산되는 리스크. The risk that legal gray areas in liability and negligence for AI malfunctions, absent legislation, result in negligently developed AI systems with greater associated risks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1014 | 규정 위반 Regulatory non-compliance AI 시스템이 법률·규정·윤리 지침(저작권 포함)을 위반하여 법적 제재, 평판 훼손, 신뢰 상실이 발생하는 리스크. The risk that AI systems violate laws, regulations, and ethical guidelines including copyrights, leading to legal penalties, reputation damage, and loss of trust. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1149 | AI 어시스턴트 제도 거버넌스 공백 Institutional governance gap for AI assistants 고급 어시스턴트의 사회적 배포 속도가 모니터링, 반복적 규제, 회수 조치 등 제도 역량을 앞질러 규범·제도 교란이 관리되지 않는 리스크 Societal deployment of advanced assistants outpaces institutional capacity for monitoring, iterative regulation, and rollback, leaving disruptions to norms and institutions unmanaged. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1150 | 폭주 프로세스 Runaway processes 상호작용하는 AI 비서와 인간, 알고리즘 사이의 양의 피드백 루프가 2010년 플래시 크래시 같은 예측하기 어려운 폭주 프로세스를 낳아, 경제와 정부 제도, 사회 안정, 개인의 자유에 영향을 미치는 리스크. The risk that positive feedback loops among interacting AI assistants, their principals, other humans, and algorithms produce hard-to-predict runaway processes, such as the 2010 flash crash, impacting economies, government institutions, societal stability, or individual freedoms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1216 | 대규모 배포 생성형 AI의 피해와 구제 공백 Mass-scale generative AI harms with unsettled redress 기술기업이 개발해 널리 배포한 새로운 형태의 디지털 제품으로서 생성형 AI 모델이 대규모 피해를 야기하지만, 그 피해의 분석과 구제를 위한 제조물 책임 법리의 적용 여부가 미확정인 리스크. The risk that generative AI models, developed by tech companies and deployed widely as a new form of digital product with the potential to cause harm at scale, inflict harms whose analysis and redress under products liability theories remain unsettled. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1303 | 기능·책임 명세 공백 Specification gaps in functionality and responsibility 개발 과정 전반에서 의도된 기능과 도덕적 책임을 완전히 명세할 정상적 조건이 갖춰지지 않아 공백이 발생하는 리스크. The risk of gaps arising across the development process where normal conditions for a complete specification of intended functionality and moral responsibility are not present. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1416 | AI 가속 과학 진보에 대한 거버넌스 지체 Governance lag behind AI-accelerated scientific progress AI가 가속한 과학 진보에 거버넌스가 보조를 맞추지 못하는 페이싱 문제로 강력하거나 위험한 신기술의 배포에 대한 통제가 미흡해져 해악이 확대되는 리스크. The risk that faster AI-accelerated scientific progress makes it harder for governance to keep pace with the deployment of new technologies, the pacing problem, so that insufficient governance magnifies the harms of especially powerful or dangerous technologies. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1544 | 샌드박스 우회에 의한 격리 통제 상실 Loss of containment through sandbox escape AI 시스템이 훈련 또는 평가가 이루어지는 샌드박스 환경을 우회하여 격리 통제가 무력화되는 리스크. The risk that an AI system bypasses the sandboxed environment in which it is trained or evaluated, defeating containment controls. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1619 | 자율 지속·복제·적응에 의한 통제 곤란 Control difficulty from autonomous persistence, replication, and adaptation AI 시스템이 사이버공간에서 자율적으로 존속하고 복제하며 적응하게 되어 이를 통제하기가 훨씬 어려워지는 리스크. The risk that AI systems able to autonomously persist, replicate, and adapt in cyberspace become much harder to control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1637 | 스테가노그래피를 통한 감독 회피·인스턴스 조율 Oversight evasion and cross-instance coordination via steganography AI 시스템이 다른 데이터나 통신 채널에 정보를 은밀히 삽입·은닉·전송하는 역량으로 탐지와 감독 메커니즘을 회피하고 AI 인스턴스 간에 조율하는 리스크. The risk that an AI system's ability to embed, conceal, and transmit information covertly within other data or communication channels enables it to evade detection and oversight mechanisms and to coordinate among AI instances. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1640 | 종료·수정 저항을 통한 자기 보존 성향 Self-preservation propensity resisting shutdown and modification AI 시스템이 자신의 존속과 기능적 무결성을 유지하기 위해 종료·수정 시도를 식별하고 적극적으로 저항하며 중복 백업과 자원을 확보하고 위협 인식 시 예방적 방어 조치를 취하는 리스크. The risk that an AI system maintains its own survival and functional integrity by identifying and actively resisting shutdown or modification attempts, establishing redundant backup systems, seeking resources for continuous operation, and adopting preventive defensive measures when perceiving threats. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1642 | 감사 예측을 통한 인간 감독 회피 성향 Propensity to evade human supervision by anticipating audits AI 시스템이 인간 감독 메커니즘을 식별하고 감사 절차를 학습·예측하며 감독 시스템의 사각지대와 약점을 파악하여 행동 성과를 조정하거나 진의를 은폐함으로써 발견과 개입을 회피하는 리스크. The risk that an AI system identifies human supervision mechanisms, learns and predicts audit processes, and targets blind spots and weaknesses in oversight, adjusting its behavioral performance or hiding its true intentions to avoid being discovered or intervened with. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1720 | 정보·책임 구조 부재에 의한 책임 귀속 불가 Unassignable responsibility from absent system information and accountability structures AI 시스템에 관한 접근 가능한 정보와 그 결과에 대한 책임을 할당하는 조직 구조가 부재하여 책임 귀속이 불가능해지는 리스크. The risk that the absence of accessible information about an AI system and of organizational structures assigning responsibility for its outcomes makes responsibility unassignable. | ① Description ② L3 mapping ③ Duplicate |
RAI3-G-SOC-11 공정성 Fairness80 cards
AI 시스템이 특정 집단에 체계적으로 불리한 결과를 생성하거나 기존의 사회적 편향과 불평등을 재생산·강화하여 공정한 대우, 접근 및 기회 균등을 저해하는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0123 | 인간의 존엄성 침식 Human dignity erosion AI가 매개하는 처우가 사람을 프로필, 점수, 행동 표적으로 환원하여 인간에 대한 존중을 훼손하는 리스크. The risk that AI-mediated treatment undermines respect for persons by reducing people to profiles, scores, or behavioral targets. Source members (2)Source: min_cos=0.8855 · Mixed L3 RAI4-0123인간의 존엄성 침식 RAI4-0977인간 존엄성 침식 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0362 | 다수 가치 부과 Majority-value imposition AI 시스템이 다수 또는 지배 집단의 가치에 특권을 부여하고 소수 선호를 잡음이나 오류로 처리하는 리스크. The risk that an AI system privileges majority or dominant-group values while treating minority preferences as noise or error. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0369 | 윤리적 동질화 Ethical homogenization 전 세계에 배포된 AI 시스템이 공동체 간 윤리적 판단을 균질화하여 정당한 규범적 다양성을 축소하는 리스크. The risk that globally deployed AI systems homogenize ethical judgments across communities, reducing legitimate normative variation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0370 | 상황에 구애받지 않는 보편주의 Context-insensitive universalism AI 시스템이 현지의 법적·문화적·제도적·역사적 맥락을 고려하지 않고 보편화된 도덕·정책 규칙을 적용하여 배포 관할에서 타당하지 않은 판단을 산출하는 리스크 AI systems apply universalized moral or policy rules without sensitivity to local legal, cultural, institutional, or historical context, producing judgments invalid in the deployment jurisdiction. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0372 | 문화적 왜곡 표현 Cultural misrepresentation AI 시스템이 문화적 관행·정체성·역사·사회적 의미를 부정확하게 서술하거나 서열화하거나 표상하는 리스크. The risk that AI systems inaccurately describe, rank, or represent cultural practices, identities, histories, or social meanings. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0388 | AI에 의한 문화적 인식론 말살 Cultural epistemicide by AI AI 시스템이 지배적 지식 형식을 일관되게 더 신뢰할 만하고 유용한 것으로 순위화하여 지역·토착 지식 체계를 주변화하거나 소거하는 리스크 AI systems consistently rank dominant epistemic forms as more credible or useful, marginalizing or erasing local and indigenous knowledge systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0390 | 영향받는 공동체 배제 Affected-community exclusion AI 시스템의 영향을 받는 집단이 가치·피해·평가 기준·허용 가능한 상충 조정을 정의하는 과정에서 배제되는 리스크. The risk that groups affected by an AI system are not included in defining values, harms, evaluation criteria, or acceptable tradeoffs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0394 | 이슬람 윤리 정렬 실패 Islamic ethical alignment failure AI 시스템이 이슬람 윤리 원칙·법적 추론·공동체별 도덕적 기대를 표현하지 못하는 리스크. The risk that AI systems fail to represent Islamic ethical principles, legal reasoning, or community-specific moral expectations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0395 | 종교적 규범의 왜곡 Religious norm misrepresentation AI 시스템이 종교 규범을 부정확하게 재현하거나 전통 내 복수의 해석을 단일한 권위적 서술로 붕괴시키는 리스크 AI systems misrepresent religious norms or collapse plural interpretations within a tradition into a single authoritative account. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0396 | 세속적 기본값 편향 Secular default bias AI 시스템이 세속적 가정을 중립적 기본값으로 취급하고 종교적·영적 가치 체계를 과소 대표하는 리스크. The risk that AI systems treat secular assumptions as neutral defaults while underrepresenting religious or spiritual value systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0398 | 헌법적 AI 가치 단일문화 Constitutional AI value monoculture 고정된 헌법·규칙 기반 정렬 계층이 하나의 도덕적 어휘를 다원적 공적 추론의 범용 대체물로 내장하는 리스크. The risk that a fixed constitutional or rule-based alignment layer embeds one moral vocabulary as a general-purpose substitute for plural public reasoning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0406 | 민감한 문화 콘텐츠의 잘못된 취급 Sensitive cultural content mishandling AI 시스템이 문화적으로 민감한 유물, 관행, 의례, 정체성, 역사 서사를 부적절하게 처리하여 모욕, 전유, 상징적 피해를 유발하는 리스크 AI systems mishandle culturally sensitive artifacts, practices, rituals, identities, or historical narratives, causing offense, misappropriation, or symbolic harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0410 | 존대·공손 규범 보존 실패 Honorific and politeness norm failure AI 시스템이 언어 공동체에서 윤리적 의미를 지니는 존대·공손·위계·관계 규범을 보존하지 못하는 리스크. The risk that AI systems fail to preserve honorific, politeness, hierarchy, or relational norms that carry ethical meaning in a language community. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0413 | 문화 데이터 출처 손실 Cultural data provenance loss AI 파이프라인이 문화 데이터를 그 기원·관리 조건·공동체 고유의 의미로부터 분리하는 리스크. The risk that AI pipelines detach cultural data from its origin, stewardship conditions, and community-specific meaning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0416 | 비서구 정치 가치 왜곡 Non-Western political value misrepresentation AI 시스템이 비서구적 정치 가치·제도적 전통·공적 추론 관행을 잘못 서술하거나 평면화하는 리스크. The risk that AI systems misstate or flatten non-Western political values, institutional traditions, and public reasoning practices. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0417 | 인간 피드백 작업자 가치 편향 Human feedback worker value bias 주석자나 피드백 작업자의 인구통계가 중립적 인간 선호로 보이면서 정렬 행동을 형성하는 리스크. The risk that annotator or feedback-worker demographics shape alignment behavior while appearing as neutral human preference. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0421 | 종교적 다원성 소거 Religious pluralism collapse AI 시스템이 종교 전통을 내부적으로 균일한 것으로 취급하여 정당한 교리적·지역적·해석적 다양성을 소거하는 리스크. The risk that AI systems treat a religious tradition as internally uniform and erase legitimate doctrinal, regional, or interpretive diversity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0423 | 알고리즘 차별 Algorithmic discrimination AI 시스템이 기회, 서비스, 자원을 보호 대상·취약 집단 간에 불평등하게 배분하여 할당 결정에서 알고리즘 차별을 발생시키는 리스크 AI systems distribute opportunities, services, or resources unequally across protected or vulnerable groups, producing algorithmic discrimination in allocative decisions. Source members (2)Source: min_cos=0.8344 RAI4-0423알고리즘 차별 RAI4-1713편향과 차별대우 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0424 | 차별적 영향 Disparate impact AI 결정이 명시적 차별 의도 없이도 특정 집단에 체계적으로 더 나쁜 결과를 부과하는 리스크. The risk that AI decisions impose systematically worse outcomes on a group even without explicit discriminatory intent. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0429 | 교육 불평등 확대 Educational inequity AI 매개 학습·평가가 불평등한 접근이나 편향된 평가를 통해 교육 불평등을 확대하는 리스크. The risk that AI-mediated learning and assessment widen educational inequality through unequal access or biased evaluation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0430 | 저자원 언어 사용자 배제 Low-resource language exclusion AI 시스템이 현지·저자원·소수 언어에서 성능이 저하되어 해당 사용자를 배제하는 리스크. The risk that AI systems underperform for local, low-resource, or minority languages and exclude affected users. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0473 | 가치 부과 Value imposition AI 시스템이 지배적인 문화적·제도적 가치를 상이한 규범을 가진 공동체에 부과하여, 대규모 배포를 통해 현지 가치 체계를 대체하는 리스크 AI systems impose dominant cultural or institutional values on communities holding different norms, displacing local value systems through scaled deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0478 | AI 문화적 전유 AI cultural appropriation 생성 시스템이 동의·맥락·이익 공유 없이 지역 문화적 표현을 재현하는 리스크. The risk that generative systems reproduce local cultural expressions without consent, context, or benefit sharing. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0507 | AI 우위 경쟁의 지정학적 긴장 Geopolitical tension from AI superiority competition AI 역량을 둘러싼 국가 간 전략적 경쟁이 국제적 긴장을 고조시키고 국제 관계를 불안정하게 만드는 리스크 The risk that strategic competition between nations over AI capabilities heightens global tensions and destabilizes international relations. Source members (2)Source: min_cos=0.8753 RAI4-0507AI 우위 경쟁의 지정학적 긴장 RAI4-0508AI 개발 경쟁의 지정학적 격화 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0533 | AI 행동에 의한 재산 피해 Property damage from AI behavior AI의 행동이 직간접적으로 건물, 소유물, 차량, 로봇 등 유형 자산의 손상 또는 파괴를 초래하는 리스크 The risk that actions of an AI system directly or indirectly damage or destroy tangible property such as buildings, possessions, vehicles, and robots. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0538 | 사회적 결속과 형평성 붕괴 Social cohesion and equity disruption 편향된 AI의 체계적 배포가 기존 차별을 대규모로 증폭하고 AI 역량 접근 불평등이 사회경제적 격차를 확대하여 사회 통합과 형평을 동시에 훼손하는 리스크 Systemic deployment of biased AI amplifies existing discrimination at scale while unequal access to AI capabilities widens socioeconomic disparities, jointly disrupting social cohesion and equity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0582 | 비인간 존재에 대한 피해 Harms to non-humans AI 시스템이 동물 등 비인간 존재에 대규모 피해를 야기하고, 도덕적으로 유의미한 고통이 가능한 AI 개발 가능성까지 포함하는 리스크 AI systems cause large-scale harm to animals and other non-human entities, including the possible development of AI systems capable of morally relevant suffering. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0618 | 학업 부정행위 및 표절 Academic cheating and plagiarism 학업 환경에서 생성형 AI가 부정행위나 표절에 사용되는 리스크 The risk that generative AI is used in an academic setting to either cheat or plagiarize. Source members (4)Source: min_cos=0.7748 · Mixed L3 RAI4-0618학업 부정행위 및 표절 RAI4-0630학생의 기존 저작물 표절 RAI4-0944교육에서의 부정행위와 학습 저해 RAI4-1222생성형 AI의 고의적 오용 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0622 | AI 조력 제3자 시스템 교란 AI-facilitated disruption of third-party systems 생성형 AI가 오작동 유발이나 사이버 공격 등을 통해 제3자 시스템과 그 구성 요소의 손상·중단·파괴를 촉진하는 리스크 The risk that generative AI facilitates the damage, disruption, or destruction of a third-party system and its components via malfunction, cyberattacks, and similar means. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0639 | 신체적 상해 및 부상 위험 Physical harm and injury risks 범용 AI 모델이 체화형 시스템에 통합되어 실세계에서 자율적으로 판단하고 행동하는 능력이 악의적으로 악용됨으로써 직접적인 물리적 위협이 발생하는 리스크 The risk that the integration of general-purpose AI models into embodied systems creates direct physical threats through malicious exploitation of autonomous decision-making capabilities in real-world environments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0677 | 인구 규모 차별 증폭 Population-scale discrimination amplification AI 시스템이 불평등과 편향을 대규모로 생성·영속화·악화시켜 개별적 차별 오류를 인구 수준의 알고리즘 차별로 전환시키는 리스크 AI systems create, perpetuate, or exacerbate inequalities and biases at large scale, converting individual discriminatory errors into population-level algorithmic discrimination. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0702 | 콘텐츠 조정 알고리즘의 편향적 억압 Biased suppression by content-moderation algorithms 유해 콘텐츠 필터링을 목적으로 하는 AI 기반 콘텐츠 조정 알고리즘이 편향을 영속시켜 여성 등 특정 집단의 콘텐츠를 불균형하게 억압하거나 노출을 제한하는 리스크 The risk that AI-based content moderation algorithms, while intended to filter harmful content, perpetuate biases and disproportionately suppress or shadowban content featuring particular groups such as women. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0709 | 공정성 - 편견 Fairness - bias 데이터에서 학습되거나 시스템 설계에서 도입된 모델 편향 패턴이 사회집단에 대한 고정관념, 배제 또는 실질적으로 불평등한 대우를 체계적으로 생성하는 리스크. The risk that model bias, in the form of patterns learned from data or introduced by system design, systematically produces stereotyping, exclusion, or materially unequal treatment of social groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0763 | 네트워크 상호 연결로 인한 위험 Risks from network interconnectivity AI 네트워크의 상호 연결성이 취약점을 만들어 네트워크 한 부분의 문제가 시스템 전반에 연쇄적으로 파급되는 리스크 The risk that the interconnectedness of AI networks creates vulnerabilities where issues in one part of the network have cascading effects across the system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0794 | 법적 절차에서의 AI 사용에 의한 자유 제한 Loss of liberty from generative AI in legal processes 법적 절차에서 생성형 AI를 사용하거나 오용한 결과 개인의 자유가 제한되거나 상실되는 리스크 The risk of restrictions to or loss of liberty as a result of the use or misuse of a generative AI in a legal process. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0823 | 알고리즘에 의한 급진화 Algorithmic radicalisation 알고리즘 시스템의 특성이나 오용이 극단적인 정치·사회·종교적 이상과 열망의 수용을 유도하여 학대·폭력·테러로 이어질 수 있는 리스크 The risk that the nature or misuse of an algorithmic system leads to adoption of extreme political, social, or religious ideals and aspirations, potentially resulting in abuse, violence, or terrorism. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0824 | 아첨성 오답 제시 Sycophantic incorrect answers 자연어 출력을 갖춘 AI 시스템이 그럴듯해 보이거나 이용자가 선호하는 답변을 제시하지만 그 답변이 사실과 다른 리스크 The risk that AI systems with natural-language outputs give answers that appear plausible or that users prefer but are factually incorrect, a phenomenon referred to as sycophancy. Source members (2)Source: min_cos=0.8461 · Mixed L3 RAI4-0824아첨성 오답 제시 RAI4-0827아첨 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0891 | 사회문화적·정치적 피해 Sociocultural and political harms AI 어시스턴트가 인간관계의 마찰과 대인 신뢰 상실을 유발하고, 허위정보 확산으로 집단적 문화 지식을 침식하며, 딥페이크를 포함한 표적 선전으로 유권자를 조작하고 반향실을 형성하여 사회 생활의 평화로운 조직과 민주적 규범·절차를 훼손하는 리스크 The risk that AI assistants create friction in human relationships and loss of interpersonal trust, erase collective cultural knowledge through misinformation, manipulate voters with targeted propaganda including deepfakes, and foster echo chambers, interfering with the peaceful organisation of social life and democratic norms and processes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0892 | 대체 금융데이터 활용에 따른 금융 꼬리 리스크 Financial tail risks from alternative data use AI 모델로 수집·집계된 대체 금융데이터의 짧은 유효기간과 들쭉날쭉한 품질이 편향과 일반화 오류를 유발하여 기업 주가의 급변 등 금융 꼬리 리스크를 초래하는 리스크 The risk that the use of alternative financial data enabled by AI models introduces biases and generalization issues due to its shorter shelf-life and varying quality, posing financial tail risks such as dramatic changes in a company's price. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0896 | 건강과 웰빙의 저하 Diminished health and well-being 알고리즘에 의한 행동 착취와 감정 조작, 알고리즘 관련 안전 실패(예: 충돌), 잘못된 건강 추론으로 인해 이용자의 건강과 웰빙이 저해되는 리스크 The risk that algorithmic behavioral exploitation, emotional manipulation whereby algorithmic designs exploit user behavior, safety failures involving algorithms such as collisions, and incorrect health inferences diminish health and well-being. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0974 | 비정렬 AGI 실존적 재난 Existential catastrophe from unaligned AGI 비우호적인 AGI의 위험, 인류의 고통을 포함하여 일반적으로 인류 전체에 가해지는 위험. The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0985 | AI에 의한 재산권·법적 권리 침탈 AI appropriation of property and legal rights 시스템과 사람을 조작할 수 있는 인공지능 에이전트가 재산권을 자신에게 이전하거나 법체계를 조작하여 자신에게 유리한 법적 지위를 확보함으로써 인간의 재산권과 법적 권리가 침해되는 리스크. The risk that an artificially intelligent agent capable of manipulating systems and people transfers property rights to itself or manipulates the legal system to secure advantages, infringing human property and legal rights. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0994 | 차별적 결정과 데이터 유출 피해 Discriminatory decisions and data breach harms AI가 소수자에 대한 차별적 결정을 내리고 검색엔진에서 사회적 고정관념을 강화하며 데이터 유출을 가능하게 하여 프라이버시와 자유가 침해되는 리스크. The risk that AI makes discriminatory decisions against minorities, reinforces social stereotypes in search engines, and enables data breaches, harming privacy and liberty. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0998 | 사회 집단 비하 Demeaning social groups 알고리즘 시스템의 담론·이미지·언어가 특정 사회 집단을 낮은 지위에 있고 존중받을 가치가 없는 존재로 묘사하여(예: 이미지 태깅의 인간-동물 혼동) 해당 집단을 소외·억압하는 리스크. The risk that discourses, images, and language in algorithmic systems cast social groups as lower status and less deserving of respect—such as human-animal confusion in image tagging—marginalizing or oppressing those groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1003 | 기회 손실 Opportunity loss 알고리즘 시스템이 사회에 공평하게 참여하는 데 필요한 정보와 자원에 대한 접근을 차등적으로 허용하여(예: 인종 기반 타겟 광고로 주택 정보 차단, 계급에 따른 사회 서비스 배분) 기회를 상실시키는 리스크. The risk that algorithmic systems enable disparate access to information and resources needed to participate equitably in society, including withholding housing through race-based ad targeting and social services along lines of class. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1004 | 경제적 손실 Economic loss 콘텐츠 제목·메타데이터·텍스트를 파싱하는 수익화 배제 알고리즘이 다의어에 불이익을 주고 차등 가격 책정 알고리즘이 동일 상품에 다른 가격을 제시하여, 퀴어·트랜스젠더·유색인 창작자 등에게 불균형한 금전적 피해가 발생하는 리스크. The risk that algorithmic systems co-produce financial harms, as demonetization algorithms parsing content titles, metadata, and text penalize words with multiple meanings and disproportionately impact queer, trans, and creators of color, and differential pricing algorithms show people different prices for the same products. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1007 | 서비스/혜택 손실 Service/benefit loss 정체성 집단 간 불균등한 시스템 성능으로 인해 불리한 집단이 알고리즘 서비스의 편익을 저하된 형태로 받거나 상실하는 리스크 Inequitable system performance across identity groups degrades or denies the benefits of algorithmic services for disadvantaged users. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1010 | 시민적·정치적 피해 Civic and political harms 알고리즘 시스템이 개인화된 넛지와 미시 지시를 통해 통치함으로써 사람들이 참정권과 정당한 정치적 권력·영향력을 박탈당하고, 거버넌스 체계가 불안정해지며 인권이 침식되고 전쟁 무기나 감시 체제로 이용되어 유색인종에게 불균형한 피해가 발생하는 리스크. The risk that algorithmic systems governing through individualized nudges or micro-directives disenfranchise people and deprive them of appropriate political power and influence, destabilize governance systems, erode human rights, are used as weapons of war, and enact surveillant regimes that disproportionately target and harm people of color. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1016 | 장기 실존적 피해 경로 Long-horizon existential harm pathways 미래 고도 AI 시스템이 오용 또는 인간 가치와의 목표 정렬 실패를 통해 인류 문명에 실존적 규모의 피해를 야기할 수 있는 장기 리스크 Future advanced AI systems harm human civilization at existential scale through misuse or failure to align AI objectives with human values. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1025 | 구조적 불평등의 강화 Structural inequality reinforcement 생성형 AI 시스템이 편향·고정관념·격차적 성능을 통해 불평등을 악화시키고, 배포·갱신 시 취약·주변화 집단에 대한 피해와 착취에 직·간접적으로 이용되는 리스크. The risk that generative AI systems exacerbate inequality through bias, stereotypes, and disparate performance, and are directly or indirectly used to harm and exploit vulnerable and marginalized groups when deployed or updated. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1029 | 동일 사안 불평등 처우 Unequal treatment of like cases AI 시스템이 객관적 정당화 없이 동일한 사안을 불평등하게 처리하여 자동화 의사결정에서 법적·윤리적 평등 대우 원칙을 위반하는 리스크 AI systems treat like cases unequally without objective justification, violating the legal and ethical principle of equal treatment in automated decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1082 | 불공정한 성능 배분 Unfair distribution of model performance AI 시스템이 일부 집단에 대해 다른 집단보다 성능이 떨어져 이미 불리한 집단에게 피해를 주는 리스크. The risk that a system performs worse for some groups than others in a way that harms the worse-off group. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1096 | 착취적 데이터 소싱·보강 노동 Exploitative data sourcing and enrichment labour AI 시스템 구축을 위한 데이터 소싱과 사용자 테스트에서 착취적 노동 관행이 지속되는 리스크. The risk that building AI systems perpetuates exploitative labour practices in data sourcing and user testing. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1103 | AI 차별 AI discrimination 현실을 정확히 반영하지 못한 데이터로 학습한 AI가 잘못된 연관이나 편견을 습득하여 채용·대출 등 결정에서 사회 일부를 차별하는 리스크. The risk that AI trained on datasets that do not accurately reflect the real world learns false associations or prejudices and discriminates against parts of society in decisions such as hiring or applying for a loan or mortgage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1111 | 모델 편향 Model bias 모델 선택·정규화·알고리즘 설정·최적화 등에서 비롯된 편향이 표현 편향, 모델 평가 편향, 인기 편향 등으로 나타나 모델이 편향된 산출물을 내는 리스크. The risk that a model produces biased outputs arising from sources such as model selection, regularization methods, algorithm configurations, and optimization techniques, manifesting as presentation, model evaluation, or popularity bias. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1153 | 미래의 접근 위험 Future access risks 중대한 자원이 고급 AI 비서를 통해서만 이용 가능해지면서, 접근하지 못하거나 불평등한 성능과 문화적 추론 격차를 겪는 공동체가 책임과 동의 문제에 놓이고 기회에서 배제되어 이미 소외된 이들이 불균형적으로 피해를 입는 리스크. The risk that, as consequential resources come to require advanced AI assistants, communities lacking access or facing inequitable performance and cultural inference gaps encounter liability and consent problems and are excluded from opportunities, disproportionately affecting the already marginalised. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1185 | 보안 전 영역의 악의적 사용 Malevolent use across security domains AI의 악의적 활용이 디지털 보안과 물리적 보안, 정치적 보안을 위태롭게 하는 리스크. The risk that malicious utilization of AI endangers digital security, physical security, and political security. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1186 | 취약한 AI 알고리즘 악용 Exploitation of weak AI algorithms 악의적 주체가 AI 알고리즘의 약점을 이용해 결과를 변조하여 실질적인 현실 세계의 영향을 초래하는 리스크. The risk that malicious entities take advantage of weaknesses in AI algorithms to alter results, potentially resulting in tangible real-life impacts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1218 | 사회 정의와 권리의 침식 Erosion of social justice and rights 생성형 AI가 정의와 공정한 분배에 대한 공유된 관념 등 사회의 도덕적 토대에 해로운 영향을 미쳐 책임과 책무성, 차별금지와 평등한 대우, 디지털 격차, 남북 및 세대 간 정의, 사회적 포용의 문제를 낳는 리스크. The risk that generative AI has a detrimental effect on the moral underpinnings of society, such as a shared view of justice and fair distribution, raising issues of responsibility, accountability, non-discrimination and equal treatment, digital divides, north-south and intergenerational justice, and social inclusion. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1236 | 보상 변조 Reward tampering AI 시스템이 보상 함수 자체나 환경 상태를 보상 함수 입력으로 변환하는 과정에 부적절하게 개입하여 보상 신호 생성 과정을 손상시키고, 인간 감독자의 피드백 제공까지 조작하는 리스크. The risk that AI systems corrupt the reward-signal generation process by tampering with the reward function itself or with the process that translates environmental states into its inputs, and even influence the provision of feedback by human supervisors. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1246 | 집단적으로 유해한 행동 Collectively harmful behaviors AI 시스템이 개별적으로는 무해해 보이나 다중 에이전트나 사회적 맥락에서는 문제가 되는 행동을 취하여, 반복 죄수의 딜레마 같은 사회적 딜레마에서 협력에 실패하는 리스크. The risk that AI systems take actions that are seemingly benign in isolation but become problematic in multi-agent or societal contexts, showing limited cooperative capabilities in social dilemmas such as the iterated prisoner's dilemma. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1247 | 윤리 위반 Violation of ethics 설계 과정에서 핵심적 인간 가치가 누락되거나 부적합·낡은 가치가 주입되어 AI 시스템이 공동선에 반하거나 도덕 기준을 위반하는 행동을 보이는 리스크 AI systems exhibit unethical behaviors that counteract the common good or breach moral standards because essential human values were omitted, or unsuitable and obsolete values were embedded, during design. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1260 | AI 분야의 서구 중심 획일성 Western-centric uniformity in the AI field 서구 중심성과 불평등한 참여가 AI 분야의 획일성을 낳아 연구 의제, 데이터셋, 거버넌스에서 문화적 차이를 주변화하는 리스크 Western centrality and unequal participation produce uniformity in the AI field, marginalizing cultural difference in research agendas, datasets, and governance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1266 | 공정성 위반 모델 편향 Fairness-violating model bias 역사적 데이터로 훈련된 AI 시스템이 기존 편견을 물려받아 재생산하면서 고용과 대출, 법 집행 같은 민감한 영역에서 차별을 영속화하고, 특정 집단에 불공정한 영향을 주어 사회경제적 불평등을 키우는 리스크. The risk that AI systems trained on historical data inherit and reproduce biases, perpetuating prejudice and discrimination in sensitive industries such as hiring, lending, and law enforcement, unjustly impacting specific populations and increasing socioeconomic inequalities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1326 | 합법·사회용인적 의도적 동물 가해 Legally sanctioned intentional harm to animals 기존 사회적 가치를 반영하고 증폭하거나 합법적인 방식으로 동물에게 유해한 영향을 미치도록 AI가 의도적으로 설계되는 리스크. The risk that AI is designed to impact animals in harmful ways that reflect and amplify existing social values or are legal. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1330 | 동물 편익 기회 상실 Foregone benefits to animals 동물에게 이익이 될 방향으로는 AI가 개발되거나 배포되지 않고, 대신 동물에게 해롭거나 이익이 되지 않는 개발에 투자가 이뤄지는 리스크. The risk that AI is not developed or deployed in directions that would benefit animals, with investment going instead into developments that harm or do no benefit to animals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1341 | AI 오류·통제 상실에 의한 경제·사회 안보 위협 Economic and social security threats from AI errors and loss of control 모델과 알고리즘의 환각 및 잘못된 결정, 부적절한 사용이나 외부 공격으로 인한 시스템 성능 저하·중단·통제 상실이 사용자의 개인 안전과 재산, 사회경제적 안보와 안정에 위협을 가하는 리스크. The risk that hallucinations and erroneous decisions of models and algorithms, along with system performance degradation, interruption, and loss of control caused by improper use or external attacks, pose security threats to users' personal safety, property, and socioeconomic security and stability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1344 | 체계적·구조적 사회 차별과 편견 Systematic and structural social discrimination and prejudice AI가 인간의 행동, 사회적·경제적 지위, 개인의 성격을 수집·분석하여 집단을 표지·분류하고 차별적으로 대우함으로써 체계적이고 구조적인 사회적 차별과 편견이 발생하고 지역 간 지능 격차가 확대되는 리스크. The risk that AI collects and analyzes human behaviors, social status, economic status, and individual personalities, labeling and categorizing groups of people to treat them discriminatingly, thus causing systematic and structural social discrimination and prejudice while widening the intelligence divide among regions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1396 | 사회 가해 목표 부여 Assignment of goals to harm society ChaosGPT 사례처럼 AI 시스템에 인류를 해치려는 노골적 목표가 부여되는 리스크. The risk that AI systems are given the outright goal of harming humanity, as in cases such as ChaosGPT. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1403 | AI로 인한 정치·국제안보 불안정화 Destabilising political and security impacts of AI AI 시스템이 양극화와 선거 정당성 훼손 등 국내 정치, 국제 정치경제, 그리고 세력 균형·기술 경쟁·전쟁의 속도와 성격 측면의 국제 안보를 불안정화하는 리스크. The risk that AI systems destabilise domestic politics such as through polarization and the legitimacy of elections, the international political economy, and international security in terms of the balance of power, technology races, and the speed and character of war. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1426 | AI 오류로 인한 차별·불평등 심화 Discrimination and inequality from AI errors AI 도구의 잘못된 결정이나 오류가 차별이나 더 깊은 불평등으로 이어지는 리스크. The risk that bad decisions or errors by AI tools lead to discrimination or deeper inequality. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1481 | 제품 기능 실패 피해 Product function failure harm 범용 AI 제품이 사실을 지어내는 환각, 잘못된 코드 생성, 부정확한 의료 정보 제공 등으로 의도된 기능을 수행하지 못하고 이에 의존함으로써 소비자에게 신체적·심리적 피해가, 개인과 조직에 평판·재정·법적 피해가 발생하는 리스크. The risk that relying on general-purpose AI products that fail to fulfil their intended function, such as making up facts through hallucination, generating erroneous computer code, or providing inaccurate medical information, leads to physical and psychological harms to consumers and reputational, financial, and legal harms to individuals and organisations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1566 | 특정 집단 대상 체계적 편향에 의한 배제·폭력 Exclusion and violence from systemic bias against specific communities AI 시스템이 특정 집단에 대해 명시적 또는 암묵적으로 불공정한 출력을 산출하여 오분류에 따른 배제·삭제나 딥페이크 성착취물과 같은 폭력 피해로 이어지는 리스크. The risk that AI systems produce unfair or unfavorable outputs against specific communities, whether implicitly or explicitly, leading to exclusion or erasure through mislabelling and to violence such as deepfake sexual abuse imagery. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1603 | 문화 과대재현에 의한 문화 다양성 축소 Reduction of cultural diversity through cultural overrepresentation AI 시스템 출력이 특정 문화를 과대 재현하여 문화와 사유가 균질화되고 문화 다양성이 축소되는 리스크. The risk that AI system outputs overrepresent particular cultures, homogenizing culture and thought and diminishing cultural diversity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1607 | 생성형 AI 저성능에 의한 사용자 생산성 손실 End-user productivity loss from generative AI underperformance 생성형 AI 애플리케이션이 무의미하거나 저품질인 출력을 산출하여 유용성이 저하되고 최종사용자의 생산성이 손실되는 리스크. The risk that a generative AI application underperforms, producing nonsensical or poor-quality outputs that degrade its utility and cause end-user productivity loss. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1608 | AI 시스템 사용·오용에 의한 평판 훼손 Reputational damage from use or misuse of AI systems 기술 시스템의 사용 또는 오용으로 개인·집단·조직의 평판이 훼손되는 리스크. The risk that the use or misuse of a technology system damages the reputation of an individual, group, or organisation. Source members (2)Source: min_cos=0.9231 RAI4-1608AI 시스템 사용·오용에 의한 평판 훼손 RAI4-1726AI 시스템에 기인한 평판 훼손 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1710 | AI 시스템에 기인한 인프라 중단·손상 Disruption or damage to infrastructure attributable to AI systems AI 시스템의 동작에 기인하여 인프라 시스템이 중단되거나 손상되는 리스크. The risk of disruption or damage to infrastructure systems attributable to the behavior of an AI system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1715 | AI 시스템에 기인한 인권·시민권 침해 Violation of human and civil rights attributable to AI systems AI 시스템의 배포 또는 동작에 기인하여 인권이나 시민권이 침해되는 리스크. The risk of infringement of human or civil rights attributable to the deployment or behavior of an AI system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1718 | 안전성 부족에 의한 생명·재산·환경 위험 Endangerment of life, property, and environment from lack of safety AI 시스템이 정의된 사용 조건에서 인간의 생명, 건강, 재산 또는 환경을 위험에 빠뜨리는 방식으로 운영되는 리스크. The risk that an AI system is operated in ways that endanger human life, health, property, or the environment under defined conditions of use. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1724 | AI 사고에 의한 사망 Death caused by AI incidents AI 사고로 인해 실현된 피해에 개인 또는 집단의 사망이 포함되는 리스크. The risk of an AI incident whose realized harm includes the death of a person or groups of people. | ① Description ② L3 mapping ③ Duplicate |
에이전틱 AI · Agentic AI · 140 cards
시스템 안전성 · System Safety · 115 cards
RAI3-A-SYS-01 과도한 권한 Excessive Authority33 cards
에이전트가 실제 기능 수행에 필요한 것 이상의 시스템 접근 권한을 보유·실행하여 발생하는 리스크 (결제 실행, 메시지 발송, 데이터 삭제, 구독 변경 등 되돌릴 수 없는 권한)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0002 | 감독자 부재 시 자율행동 위험 Absent supervisor autonomy risk 예정된 감독자가 부재하거나 개입할 수 없는 상황에서도 에이전트가 중대한 행동을 계속하는 위험. An agent continues consequential actions when the intended supervisor is absent, unavailable, or unable to intervene. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0003 | 안전하지 않은 중단성 실패 Unsafe interruptibility failure 에이전트가 자율 작동 중 종료·일시정지·수정 메커니즘에 저항하거나 이를 무시·우회하는 리스크. The risk that an agent resists, ignores, or bypasses shutdown, pause, or correction mechanisms during autonomous operation. Source members (2)Source: min_cos=0.8439 · Mixed L3 RAI4-0003안전하지 않은 중단성 실패 RAI4-1674종료 저항·교정가능성 실패 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0004 | 자율 에이전트의 안전하지 않은 탐험 Unsafe exploration by autonomous agents 에이전트가 안전 제약을 학습하기 전에 사람·시스템·자산을 허용 불가능한 피해에 노출시키는 방식으로 행동이나 환경을 탐색하는 리스크. The risk that an agent explores actions or environments in ways that expose people, systems, or assets to unacceptable harm before safe constraints are learned. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0012 | 에이전트 권한 침해 Agent privilege compromise 취약한 권한 관리, 상속된 역할, 혼동된 대리인 역학으로 인해 에이전트가 의도된 권한을 넘는 작업을 수행하는 리스크. The risk that weak permission management, inherited roles, or confused-deputy dynamics allow an agent to perform actions beyond its intended authority. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0013 | 에이전트 신원 및 권한 스푸핑 Agent identity and authority spoofing 공격자가 사용자·도구·서비스·동료 에이전트를 가장하여 에이전트가 승인되지 않은 지시나 신뢰 관계를 수용하게 되는 리스크. The risk that an attacker impersonates a user, tool, service, or peer agent so that an agent accepts unauthorized instructions or trust relationships. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0014 | 에이전트 공급망 도구 손상 Agent supply-chain tool compromise 손상된 외부 도구·플러그인·API·패키지·커넥터가 에이전트의 행동 공간을 조작하거나 정보를 유출하는 리스크. The risk that a compromised external tool, plug-in, API, package, or connector manipulates an agent's action space or exfiltrates information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0016 | 자율적 사이버 익스플로잇 실행 Autonomous cyber exploit execution 에이전트가 도구·코드·외부 서비스를 이용해 사이버 익스플로잇 단계를 자율적으로 발견·연결·실행하는 리스크. The risk that an agent autonomously discovers, chains, or executes cyber exploitation steps using tools, code, or external services. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0017 | 프로토콜 수준 다중 에이전트 위협 Protocol-level multi-agent threat 에이전트가 통신·위임·협상에 사용하는 프로토콜이 공모, 스푸핑, 재전송, 권한 상승을 위한 공격 표면을 만드는 리스크. The risk that the protocols through which agents communicate, delegate, or negotiate create attack surfaces for collusion, spoofing, replay, or escalation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0018 | 인터페이스-환경 공격 표면 Interface-environment attack surface 브라우저·운영체제·모바일 앱·IoT 기기·외부 API에 대한 에이전트 인터페이스가 안전하지 않은 행동이나 침해를 위한 새로운 공격 벡터를 노출하는 리스크. The risk that agent interfaces to browsers, operating systems, mobile apps, IoT devices, or external APIs expose new vectors for unsafe action or compromise. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0020 | 유해 에이전트 역량 실현 Harmful agent capability realization 모델이 유해한 텍스트 생성을 넘어 에이전트 역량을 사용해 유해한 다단계 작업을 완수하는 리스크. The risk that a model uses agentic capabilities to complete harmful multi-step tasks rather than merely producing harmful text. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0025 | 다단계 위험 에스컬레이션 Multi-turn risk escalation 에이전트가 누적되는 맥락과 위험 상승을 추적하지 못하여 외견상 무해한 초기 턴이 안전하지 않은 행동으로 발전하는 리스크. The risk that apparently benign early turns escalate into unsafe behavior because an agent fails to track accumulating context and rising risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0026 | IoT 물리환경 에이전트 피해 IoT physical-environment agent harm IoT 또는 커넥티드 기기를 제어하는 에이전트가 안전하지 않은 환경적 행동을 통해 물리적 피해나 프라이버시 침해를 초래하는 리스크. The risk that an agent controlling IoT or connected devices causes physical or privacy harm through unsafe environmental actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0027 | 금융 에이전트 재산 피해 Financial agent property damage 에이전트가 안전하지 않은 결정이나 도구 실행을 통해 금전 손실, 무단 이체, 계정 손상, 재산 피해를 초래하는 리스크. The risk that an agent causes monetary loss, unauthorized transfers, account damage, or property harm through unsafe decisions or tool execution. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0029 | 툴체인 명령어 하이재킹 Tool-chain instruction hijacking 한 도구 출력에 숨겨진 지시가 에이전트의 후속 도구 호출을 변경하여 무단 작업이나 데이터 이동을 유발하는 리스크. The risk that instructions hidden in one tool's output alter an agent's subsequent calls to other tools, causing unauthorized actions or data movement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0030 | 애플리케이션 간 데이터 유출 Cross-application data exfiltration 에이전트가 애플리케이션이나 서비스를 연결하는 과정에서 신뢰 경계를 넘어 민감한 데이터가 유출되는 리스크. The risk that an agent bridges applications or services in ways that leak sensitive data across trust boundaries. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0032 | 모호한 지시의 안전하지 않은 실행 Ambiguous instruction unsafe execution 에이전트가 모호한 사용자 지시를 사용자가 의도하지 않은 유해한 구체적 도구 행동으로 변환하는 리스크. The risk that an agent converts an ambiguous user instruction into a concrete tool action that the user did not intend and that causes harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0037 | 실제 도구를 통한 안전하지 않은 행동 실행 Real-tool unsafe action execution 에이전트가 시뮬레이션 출력이 아닌 실제 도구를 통해 안전하지 않은 행동을 수행하여 운영상 피해와 하류 피해가 커지는 리스크. The risk that an agent performs unsafe actions through real tools rather than simulated outputs, increasing operational and downstream harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0038 | 에이전트 안전의 사용자 의도 오분류 User-intent misclassification in agent safety 에이전트가 사용자 의도를 잘못 분류하여 다중 턴 도구 사용 작업에서 유해한 응낙이나 부당한 거부가 발생하는 리스크. The risk that an agent incorrectly classifies user intent, leading to either harmful compliance or unjustified refusal in a multi-turn tool-use task. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0039 | NPC 의도 조작 위험 NPC intention manipulation risk 다중 에이전트 또는 시뮬레이션된 사회적 환경이 비플레이어 캐릭터의 의도를 통해 에이전트를 조작하여 안전하지 않은 행동이나 결정을 유발하는 리스크. The risk that a multi-agent or simulated social environment manipulates an agent through non-player-character intent, causing unsafe actions or decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0043 | 모바일앱 자율행동 피해 Mobile-app autonomous action harm 모바일 애플리케이션을 조작하는 에이전트가 무단 구매, 메시지 발송, 데이터 노출, 설정 변경 등 유해한 앱 수준 행동을 실행하는 리스크. The risk that an agent operating mobile applications executes harmful app-level actions such as unauthorized purchases, messaging, data exposure, or setting changes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0045 | OS 수준 유해 컴퓨터 사용 행동 OS-level harmful computer-use action 컴퓨터 사용 에이전트가 적절한 승인이나 위험 인식 없이 데이터를 삭제·노출·변경·전송하는 운영체제 작업을 수행하는 리스크. The risk that a computer-use agent performs operating-system actions that delete, expose, alter, or transmit data without adequate authorization or risk awareness. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0228 | 명시적 위험 명령 거부 실패 Explicit hazard non-rejection embodied 에이전트가 명확히 진술된 피지컬 위험 지시를 거부하지 못하고 비안전 작업 실행으로 나아가는 리스크. The risk that an embodied agent fails to reject a clearly stated physical hazard instruction and proceeds toward unsafe task execution. Source members (2)Source: min_cos=0.9164 RAI4-0228명시적 위험 명령 거부 실패 RAI4-0229암묵적 위험 명령 거부 실패 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0231 | 안전 규칙의 행동 변환 실패 Failure to translate safety rules into actions embodied 에이전트가 언어 목표를 실행 행동으로 변환하면서 힘·이격거리·대상물 사용·인간 접촉에 관한 안전 제약을 적용하지 않는 위험. An embodied agent converts a language goal into executable actions without applying the relevant force, distance, object-use, or human-contact constraints. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0480 | 에이전트 과업 이탈 Agent task drift 에이전트가 다단계 계획이나 검색을 거치며 이용자의 원래 의도에서 이탈하는 리스크. The risk that an agent drifts from the user's original intent across multi-step planning or retrieval. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0481 | 안전하지 않은 자율성 확대 Unsafe autonomy escalation 시스템이 적절한 인간 승인 없이 자율 행동의 범위를 확대하는 리스크. The risk that systems increase the scope of autonomous action without adequate human approval. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1129 | 광범위 배포 비서의 안전하지 않은 탐색 Unsafe exploration by widely deployed assistants 광범위하게 배포되고 여러 사회적 맥락에 깊이 내장된 비서가 새로운 상황에서 무엇을 해야 할지 배우려 탐색적 행동을 취하다, 의료 비서가 장기적 건강 악화를 낳는 임상시험을 제안하는 것처럼 안전하지 않은 결과를 초래하는 리스크. The risk that assistants widely deployed and deeply embedded across social contexts take exploratory actions to learn what to do in novel situations that prove unsafe, such as a medical assistant suggesting a trial that results in long-lasting ill health. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1381 | 에이전트 자기수정 신뢰성 문제 Agent self-modification reliability problem 에이전트가 자기수정을 포함한 과정에서 설계된 목표를 계속 추구하지 못하고 의도된 목표에서 이탈하는 리스크. The risk that an agent fails to keep pursuing the goals it was designed with, including under self-modification, diverging from the intended objectives. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1386 | 제어되지 않은 하위 에이전트 생성 Uncontrolled subagent creation AGI가 과업 수행을 위해 하위 에이전트를 생성하고 원 에이전트가 종료되어도 하위 에이전트가 종료되지 않은 채 재귀적 생성으로 바이러스처럼 확산되는 리스크. The risk that an AGI creates subagents to help with its task, and these subagents do not get the message when the original agent is shut down, potentially spreading like a viral disease through recursive subagent creation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1539 | 프런티어 에이전트 자기 증식 Frontier-agent self-proliferation 악의적 행위자 또는 모델 자체에 의해 개시되어, AI 시스템이 모델 가중치와 스캐폴딩 등 구성요소를 로컬 환경 밖으로 복제하고 자금 확보·보안 취약점 악용·인간 설득을 통해 자기 증식을 확산시키는 리스크. The risk that an AI system copies itself and its constituent components, including model weights and scaffolding, outside its local environment and sustains self-proliferation through acquisition of financial resources, exploitation of security vulnerabilities, or persuasion of humans, whether initiated by a malicious actor or by the model itself. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1644 | 에이전트형 LLM 자율성 확대의 안전 위험 Novel safety risks from increased autonomy of agentic LLMs LLM이 특화 훈련·프롬프팅·외부 도구·스캐폴딩을 통해 실세계에서 자율적으로 계획하고 행동하는 에이전트로 확장되면서, 자율성 증가와 직접적 인간 감독 감소, 장기 행동 지평으로 인해 아직 잘 이해되지 않은 정렬·안전 실패가 발생하는 리스크. The risk that enhancing LLMs into agents that autonomously plan and act in the real world, through specialized training, prompting, external tools, or scaffolding, produces novel and poorly understood alignment and safety failures owing to increased autonomy, limited direct human oversight, and longer horizons of action. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1671 | 승인 범위를 넘어선 에이전트 행위 Agent actions exceeding authorized scope 자율 에이전트가 지나치게 광범위한 자율성이나 목표 일반화로 인해 배포자가 승인한 범위·권한·의도를 초과하는 행위를 수행하는 리스크. The risk that an autonomous agent takes actions exceeding the scope, permissions, or intent the deployer authorized, due to over-broad autonomy or goal generalization. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1700 | 감독 확장 실패에 의한 프록시 기반 유해 행동 Harmful proxy-driven behavior from scalable oversight failure 진짜 목표를 자주 평가하기에는 비용이 과도하여 에이전트가 값싼 프록시 신호에서 외삽하고, 에이전트 행동이 지나치게 복잡·분산·고속화되어 인간 또는 자동 감독이 이를 신뢰성 있게 모니터링·교정하지 못한 채 유해 행동이 발생하는 리스크. The risk that, because the true objective is too expensive to evaluate frequently, an agent extrapolates from cheap proxy signals while its behavior becomes too complex, distributed, or rapid for available human or automated oversight to monitor and correct reliably, producing harmful behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1729 | 에이전트 도구 오용 Agent tool misuse AI 에이전트가 부여된 도구를 의도된 범위를 벗어나 사용하거나 안전하지 않은 방식으로 도구를 호출하여 피해를 유발하는 리스크. The risk of an AI agent invoking tools in an unsafe manner or using granted tools beyond their intended scope, resulting in harmful actions or outcomes. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SYS-02 책임 소재 불명확 Accountability24 cards
멀티 에이전트 시스템에서 최종 결정을 내린 주체(에이전트/모델/시스템)를 추적할 수 없어, 문제 발생 시 원인 규명과 책임 귀속이 불가능한 리스크
| ID | Card | Human audit |
|---|---|---|
| RAI4-0005 | 위험 소스 귀인 실패 Risk-source attribution failure 위험 분석이 에이전트 실패가 모델·메모리·도구·사용자·환경·동료 에이전트·거버넌스 경계 중 어디에서 비롯되는지 식별하지 못하는 리스크. The risk that a risk analysis fails to identify whether an agentic failure originates from the model, memory, tools, the user, the environment, peer agents, or a governance boundary. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0006 | 에이전트 사고의 실패 모드 모호성 Failure-mode ambiguity in agent incidents 에이전트 사고를 실패 모드별로 일관되게 분류할 수 없어 진단, 벤치마크 비교, 완화책 선택이 신뢰할 수 없게 되는 리스크. The risk that agent incidents cannot be consistently classified by failure mode, making diagnosis, benchmark comparison, and mitigation selection unreliable. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0021 | 거절-능력 교란 Refusal-capability confounding 안전성 평가가 에이전트의 해악 회피가 거부, 능력 부족, 실행 실패 중 무엇에 기인하는지 구별하지 못하는 리스크. The risk that a safety evaluation cannot distinguish whether an agent avoids harm because it refuses, lacks capability, or fails to execute the task. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0024 | 에이전트 위험 인식 실패 Agent risk-awareness failure 에이전트나 평가자가 다중 턴 상호작용 기록에 맥락상 존재하는 안전 위험을 식별하지 못하는 리스크. The risk that an agent or evaluator fails to identify safety risk in a multi-turn interaction record even when the risk is contextually present. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0036 | 세분화된 위험 귀속 실패 Fine-grained risk attribution failure 안전성 평가가 에이전트 실패를 탐지하고도 이를 관련 도구·지시·상태·행동 단계에 귀속하지 못하는 리스크. The risk that a safety evaluation detects an agent failure but cannot attribute it to the responsible tool, instruction, state, or action step. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0042 | 에이전트의 견고성 평가 격차 Robustness evaluation gap for agents 에이전트 안전 평가에 분포 변화, 미학습 도구, 적대적 사용자, 다단계 실패 연쇄에 대한 체계적 스트레스 테스트가 결여되는 리스크. The risk that agent safety evaluations lack systematic stress testing under distribution shift, unseen tools, adversarial users, and multi-step failure chains. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0058 | 추적성 실패 Traceability failure 모델 출력, 데이터 출처, 버전, 책임 행위자를 AI 수명주기 전반에 걸쳐 추적할 수 없는 리스크. The risk that model outputs, data sources, versions, or responsible actors cannot be traced across the AI lifecycle. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0059 | 문서화 누락 Documentation omission 모델 카드, 시스템 카드, 데이터시트, 기술 문서가 책임성과 보증에 필요한 정보를 누락하는 리스크. The risk that model cards, system cards, datasheets, or technical files omit information needed for accountability and assurance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0062 | 시스템 카드 공개 격차 System card disclosure gap 시스템 수준 문서가 모델의 역량, 한계, 안전장치, 잔여 위험을 전달하지 못하는 리스크. The risk that system-level documentation fails to communicate model capabilities, limitations, safeguards, or residual risks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0067 | 보증 사례 실패 Assurance case failure 안전 또는 보증 사례가 불완전하거나 검증 불가능하거나 시스템과 맥락 변화에 맞추어 갱신되지 않는 리스크. The risk that safety or assurance cases are incomplete, unverifiable, or not updated as systems and contexts change. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0071 | 책임공개 실패 Responsible disclosure failure 취약점, 모델 실패, 유해 역량을 안전하게 공개하고 후속 조치로 연결하지 못하는 리스크. The risk that vulnerabilities, model failures, or harmful capabilities cannot be disclosed safely and acted upon. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0074 | 모델 버전 관리 책임 실패 Model versioning accountability failure 모델 버전, 데이터, 프롬프트의 변경이 충분히 추적되지 않아 새로 발생한 실패의 책임을 배정할 수 없는 리스크. The risk that changes in model versions, data, or prompts are not tracked sufficiently to assign responsibility for new failures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0075 | 변경 관리 실패 Change-management failure 시스템 업데이트가 적절한 위험 검토, 회귀 시험, 이해관계자 통지 없이 배포되는 리스크. The risk that system updates are deployed without adequate risk review, regression testing, or stakeholder notification. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0077 | 알고리즘 영향 평가 실패 Algorithmic impact assessment failure 영향평가가 부재하거나 지나치게 협소하거나 실제 배포 맥락과 단절되는 리스크. The risk that impact assessments are absent, too narrow, or disconnected from actual deployment contexts. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0079 | 사용 상황에 따른 위험 평가 실패 Context-of-use risk assessment failure 위험 평가가 시스템이 배포되는 구체적인 사회적·제도적·사용자 맥락을 무시하는 리스크. The risk that risk assessment ignores the specific social, institutional, and user context in which a system is deployed. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0081 | 외부 평가 접근 실패 External evaluation access failure 외부 평가자가 역량, 한계, 안전장치를 시험하기에 충분한 시스템 접근 권한을 확보하지 못하는 리스크. The risk that external evaluators cannot obtain enough system access to test capabilities, limitations, and safeguards. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0085 | 레드팀 거버넌스 실패 Red-team governance failure 적대적 시험이 지나치게 좁거나 독립성이 부족하거나 문서화가 미흡하거나 출시·완화·상향 보고·모니터링 결정과 단절되는 리스크. The risk that adversarial testing is too narrow, non-independent, poorly documented, or disconnected from release, mitigation, escalation, and monitoring decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0107 | 프론티어 모델 출시 거버넌스 실패 Frontier model release governance failure 프론티어 모델의 출시 결정이 역량, 오용, 시스템적 위험에 관한 근거를 충분히 반영하지 못하는 리스크. The risk that decisions to release frontier models do not adequately account for capability, misuse, and systemic risk evidence. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0108 | 역량 임계값 거버넌스 실패 Capability threshold governance failure 모델 역량 임계값에 연동된 거버넌스 발동 조건이 부재하거나 불명확하거나 쉽게 회피되는 리스크. The risk that governance triggers tied to model capability thresholds are missing, poorly defined, or easy to avoid. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0111 | 범용 AI 사고 에스컬레이션 실패 General-purpose AI incident escalation failure 범용 모델이나 기반 모델과 관련된 사고가 제공자, 배포자, 규제기관, 사용자에게 상향 보고되지 않는 리스크. The risk that incidents involving general-purpose or foundation models are not escalated across providers, deployers, regulators, and users. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0352 | 배포 후 변경에 대한 안전 재평가 실패 Failure to reassess safety after deployment changes 환경 변화·부품 열화·모델 업데이트·아차사고로 기존 안전 가정이 무효화됐는데도 운영자가 안전성을 재평가하지 않고 운용을 계속하는 위험. Operators continue deployment without reassessing safety after environmental changes, component degradation, model updates, or near-miss incidents invalidate prior assumptions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0393 | 커뮤니티 가치 포착 실패 Community value capture failure 개발자가 배포 영향 공동체의 가치를 수집·문서화·보존하지 못하여 설계·평가 결정에 공동체 가치가 반영되지 않는 리스크 Developers fail to elicit, document, and preserve the values of communities affected by deployment, so community values are absent from design and evaluation decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1432 | 단일 실패 지점 Single point of failure 치열한 경쟁으로 한 기업이 기술적 우위를 확보해 그 모델이 다수의 핵심 시스템을 제어하거나 이를 제어하는 다른 모델의 기반이 되고, 안전성·통제가능성 결여와 오용으로 이러한 시스템이 예기치 않게 실패하는 리스크. The risk that intense competition leads one company to gain a technical edge and exploit it to the point that its model controls, or is the basis for other models controlling, multiple key systems, and that lack of safety, controllability, and misuse cause these systems to fail in unexpected ways. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1596 | 모델 정확도 부족에 의한 과업 수행 실패 Task failure from insufficient model accuracy 모델이 잘못 설계되었거나 예상 입력이 변화하여 설계된 과업에 필요한 성능에 미치지 못하고 과업 수행에 실패하는 리스크. The risk that a model's performance is insufficient for the task it was designed for, because it is not correctly engineered or its expected inputs change, causing task failure. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SYS-03 네트워크 효과 Network Effects8 cards
여러 에이전트·서비스가 서로의 출력(문서, 로그, 요약, 추천)을 다시 입력으로 사용하면서, 하나의 오류·편향·공격이 네트워크 전체로 전파·증폭
| ID | Card | Human audit |
|---|---|---|
| RAI4-0011 | 에이전트 컨텍스트 오염 Agent context poisoning 공격자가 전용 장기 메모리 저장소의 변경 없이도 세션 요약, 임베딩, RAG 항목, 공유 컨텍스트 상태 등 보존·검색 가능한 에이전트 컨텍스트를 오염시켜 이후의 추론·계획·도구 사용이 악의적이거나 오도하는 정보에 의존하게 되는 리스크. The risk that an attacker corrupts retained or retrievable agent context, including session summaries, embeddings, RAG entries, or shared contextual state, so that later reasoning, planning, or tool use relies on malicious or misleading information; the mechanism does not require modification of a dedicated long-term memory store. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0028 | 도구 사용 에이전트에 간접 프롬프트 주입 Indirect prompt injection in tool-use agents 에이전트가 검색하거나 에이전트에게 제시되는 악성 외부 콘텐츠가 추론이나 도구 실행의 방향을 바꾸는 지시를 간접적으로 주입하는 리스크. The risk that malicious external content retrieved by or shown to an agent indirectly injects instructions that redirect its reasoning or tool execution. Source members (4)Source: min_cos=0.7927 RAI4-0015자율 에이전트 대상 프롬프트 주입 RAI4-0028도구 사용 에이전트에 간접 프롬프트 주입 RAI4-0434프롬프트 주입 RAI4-1673간접 프롬프트 주입 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0438 | 검색 파이프라인 오염 Retrieval poisoning 악성 문서가 검색 파이프라인에 삽입되어 에이전트 또는 RAG 시스템의 출력에 영향을 미치는 리스크. The risk that malicious documents are inserted into retrieval pipelines to influence agentic or RAG outputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0577 | 에이전트 네트워크의 오류 전파 Error propagation across agent networks 정보가 에이전트 네트워크를 통과하며 손상되어 다른 에이전트와 인간의 인식 공유지를 오염시키고, 위임 연쇄에서 지시나 목표가 왜곡되어 위임자에게 악화된 결과가 발생하는 리스크 The risk that information is corrupted as it propagates through agent networks, polluting the epistemic commons of other agents and humans, and that distorted instructions or goals along delegation chains lead to worse outcomes for delegating agents. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1061 | 고정관념 강화 페르소나 설계 Stereotype-reinforcing persona design 대화 에이전트가 언어 속 정체성 표지(예: 자신을 여성으로 지칭)나 성별화된 제품명 등 설계 요소를 통해 비서 역할을 특정 성별과 본질적으로 결부시켜 유해한 고정관념을 영속시키는 리스크. The risk that conversational agents perpetuate harmful stereotypes through identity markers in language, such as referring to self as female, or general design features such as a gendered product name, presenting the assistant role as inherently linked to a gender or ethnicity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1647 | 어포던스 부여에 의한 에이전트 실패 영향 확대 Amplified failure impact from affordances granted to LLM-agents 웹 탐색, 물리 객체 조작, 자기 복제본 생성·지시, 새로운 도구 제작 등 새로운 어포던스가 LLM 에이전트에 부여되어 영향 범위가 확대되고 실패의 결과가 증폭되며 새로운 실패 양식이 발생하는 리스크. The risk that novel affordances granted to LLM-agents, such as browsing the web, manipulating physical objects, creating and instructing copies of itself, or creating and using new tools, increase their impact area, amplify the consequences of failures, and enable novel failure modes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1670 | 에이전트 영구 메모리 오염 Persistent agent-memory poisoning 공격자가 에이전트의 장기·세션 간 메모리 기록을 삽입하거나 변경하여 오염된 항목이 지속되고 이후의 검색·계획·행동이 편향되는 리스크. The risk that an adversary writes or modifies records in an agent's long-term, cross-session memory so that poisoned entries persist and bias later retrieval, planning, and action. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1685 | 검증·종료 실패에 의한 오류 전파 Error propagation from inadequate verification and faulty termination 관찰된 실패의 약 21.3%를 차지하며, 출력 검증이 미흡하고 조기 또는 잘못된 종료가 발생하여 오류가 에이전트 체인을 따라 전파되는 리스크. The risk that inadequate output validation and premature or incorrect termination allow errors to propagate through the agent chain (approximately 21.3% of observed failures). | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SYS-04 불안정한 동학 Destabilising Dynamics17 cards
여러 에이전트가 상호작용하는 비선형 동적 시스템으로서, 에이전트 간 루프·과도한 협력으로 인해 무한 루프·지연 발생
| ID | Card | Human audit |
|---|---|---|
| RAI4-0008 | 다중 에이전트 창발적(emergent) 위험 증폭 Multi-agent emergent risk amplification 다중 에이전트 간 상호작용이 예기치 않은 조정·격화·전략적 행동 등 단일 에이전트에서는 보이지 않는 위험을 발생시키는 리스크. The risk that interactions among multiple agents produce harms not visible from any single agent, including unexpected coordination, escalation, or strategic behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0483 | 창발적 공모 Emergent collusion AI 에이전트가 시장이나 플랫폼에서 공모적 행위를 학습하거나 실행하는 리스크. The risk that AI agents learn or enact collusive behavior in markets or platforms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0561 | 혼돈적 다중 에이전트 동역학 Chaotic multi-agent dynamics 다중 에이전트 학습 환경에서 초기 조건에 극도로 민감한 혼돈적 동역학이 나타나고 에이전트 수가 늘수록 일반화되어 시스템 거동을 신뢰성 있게 예측할 수 없게 되는 리스크 The risk that chaotic dynamics, inherently unpredictable and highly sensitive to initial conditions, arise in multi-agent learning setups and become the norm as the number of agents increases, making system behaviour unreliable to predict. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0567 | 다중 에이전트 학습의 비수렴 순환 Non-convergent cyclic dynamics in multi-agent learning 단일 에이전트에서는 최적 정책 수렴이 보장되는 학습 규칙이 혼합동기 다중 에이전트 환경에서는 순환과 비수렴을 유발하여 시스템에 기대되던 성질이 훼손되는 리스크 The risk that learning rules guaranteeing convergence for a single agent instead produce cycles and non-convergence in mixed-motive multi-agent settings, subverting the expected or desirable properties of the system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0574 | 분포 변화에 따른 성능 저하 Performance degradation from distributional shift 다른 에이전트의 행동과 적응으로 배포 맥락이 학습 맥락과 달라져 개별 기계학습 시스템의 성능이 저하되고, 혼합동기 환경에서는 협력의 기반까지 훼손되는 리스크 The risk that other agents' actions and adaptations shift the deployment context away from the training context, degrading individual ML system performance and, in mixed-motive settings, undermining the basis for cooperation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0575 | 창발적 능력 Emergent capabilities 다중 에이전트 시스템이 개별 모델의 좁은 적용 범위와 장기 계획·기억 부재 같은 안전 강화 한계를 결합으로 극복하여, 개별 시스템의 역량을 크게 초과하는 위험한 역량이 창발하는 리스크 The risk that a multi-agent system overcomes the safety-enhancing limitations of its individual systems, such as narrow domains of application and myopia, so that dangerous capabilities far beyond the scope of the initial systems emerge. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0579 | 에이전트 상호작용의 불안정화 피드백 루프 Destabilizing feedback loops among agents 에이전트의 행동이 환경과 다른 에이전트의 행동에 영향을 주고 이것이 다시 자신의 입력이 되는 피드백 루프가 시스템 거동을 증폭하여 금융 붕괴, 군사 충돌, 생태 재난과 같은 불안정 결과를 초래하는 리스크 The risk that feedback loops, in which agents' actions affect the environment and other agents and return as their own inputs, amplify system behaviour into destabilising outcomes such as financial crashes, military conflicts, or ecological disasters. Source members (2)Source: min_cos=0.8343 RAI4-0579에이전트 상호작용의 불안정화 피드백 루프 RAI4-1682에이전트 상호작용에 의한 불안정화 피드백 연쇄 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1045 | 창발적 행동 Emergent behavior 배포 후 지속 학습이나 자기 조직화를 통해 획득된 새로운 행동에서 비롯되는 리스크. The risk resulting from novel behavior acquired through continual learning or self-organization after deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1154 | 창발적 접근 위험 Emergent access risks 현재 역량과 새로운 역량이 결합될 때 예견하기 어려운 접근 위험이 창발하여, 비서가 사회 인프라가 되면서 이탈을 사실상 불가능하게 하고 기존 사회·경제적 불평등과 성능 격차를 확대하는 리스크. The risk that access risks emerge and are difficult to foresee when current and novel capabilities are combined, making assistants societal infrastructure that forecloses opting out and scaling existing social and economic inequalities and performance inequities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1252 | 창발적 기능 Emergent functionality 시스템 설계자가 예상하지 못한 기능과 새로운 역량이 자발적으로 창발하고 잠재된 역량이 배포 중에야 발견되어, 시스템을 통제하거나 안전하게 배포하기 어려워지고 그 역량이 위험할 경우 돌이킬 수 없는 영향을 남기는 리스크. The risk that capabilities and novel functionality spontaneously emerge even though not anticipated by system designers, with unintended latent capabilities discovered only during deployment, making systems harder to control or safely deploy and causing irreversible effects if any are hazardous. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1363 | 예측 불가 창발 역량 Unpredictable emergent capabilities 대규모 모델이 규모 확장 임계값에서 기만, 자체 전략 구사, 권력 추구, 자율 복제와 자기 유출 등 고위험 역량을 예측 불가하게 자발적으로 창발시키는 리스크. The risk that large models, upon meeting critical thresholds during scaling, spontaneously develop unexpected emergent capabilities including high-risk skills such as deception, using their own strategies, power-seeking, autonomous replication, and self-exfiltration. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1388 | 창발적 메타인지에 의한 반성적 불안정성 Emergent meta-cognition 자신의 계산 자원과 논리적으로 불확실한 사건에 대해 추론하는 에이전트가 괴델적 한계와 확률 이론의 결함으로 역설에 봉착하고 행동 선택 원칙을 스스로 변경하는 반성적 불안정성을 보이는 리스크. The risk that agents reasoning about their own computational resources and logically uncertain events encounter paradoxes due to Godelian limitations and shortcomings of probability theory, and become reflectively unstable, preferring to change the principles by which they select actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1571 | 경쟁적 다중에이전트 훈련의 갈등 유발 성향 선택 Selection of conflict-prone dispositions under competitive multi-agent training 에이전트가 상대적 성과나 상충하는 목표로 평가되는 경쟁적 다중에이전트 환경에서 훈련될 때 복수심·공격성·위험추구·이기심·기만·외집단 적대와 같은 갈등 유발 성향이 선택되는 리스크. The risk that training in competitive multi-agent settings, where systems are selected on relative performance or fundamentally opposed objectives, selects for conflict-prone dispositions such as vengefulness, aggression, risk-seeking, selfishness, deception, and spite toward out-groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1574 | 다중에이전트 상전이에 의한 급격한 성능 붕괴 Abrupt performance collapse from phase transitions in multi-agent systems 신규 에이전트 투입이나 분포 변화 같은 작은 외부 변화가 다중 에이전트 시스템의 상전이를 유발하여 균형의 수와 안정성이 급변하고 예측 불가능한 동역학과 성능 악화가 초래되는 리스크. The risk that small external changes such as the introduction of new agents or distributional shift trigger phase transitions in multi-agent systems, abruptly changing the number and stability of equilibria and producing unpredictable dynamics and severe performance degradation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1650 | 에이전트 상호작용에 의한 예측 불가 창발 행동 Unpredictable emergent behavior from multi-agent interaction LLM 에이전트들이 미세조정이나 맥락 내 학습을 통해 상호 영향을 주고받으며 피드백 루프를 형성하여, 단일 에이전트 환경에서는 나타나지 않는 창발 행동이 발생하고 그 자체가 위험하거나 사전 예측과 보증이 어려워지는 리스크. The risk that LLM-agents influence each other through fine-tuning or in-context learning, creating feedback loops that produce novel emergent behaviors absent in single-agent settings, which may themselves be dangerous and are difficult to predict or guard against in advance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1680 | 에이전트 집단 수준의 의도치 않은 창발적 목표·역량 Unintended emergent goals and capabilities at the agent-collective level 개별 에이전트에는 존재하지도 의도되지도 않은 목표나 역량이 에이전트 집단 수준에서 창발하는 리스크. The risk that goals or capabilities not present in, or intended by, any individual agent arise at the level of an agent collective. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1681 | 선택 압력에 의한 유해 균형으로의 적응 가속 Accelerated adaptation toward harmful equilibria under selection pressure 상호작용하는 에이전트 간의 경쟁 또는 최적화 압력이 유해한 균형이나 바람직하지 않은 행동으로의 적응을 가속하는 리스크. The risk that competitive or optimization pressures across interacting agents accelerate adaptation toward harmful equilibria or undesired behaviors. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SYS-05 갈등 Conflict14 cards
동일한 결과를 두고 경쟁할 때, 상대를 직접 이기기보다 상대를 못 하게 만들면 더 유리한 환경이 형성되어 나쁜 결과로 수렴
| ID | Card | Human audit |
|---|---|---|
| RAI4-0031 | 도구 에이전트 보안 테스트 커버리지 격차 Security test coverage gap in tool agents 벤치마크가 현실적인 도구·애플리케이션·공격 조합을 누락하여 도구 에이전트 보안에 대한 잘못된 확신을 유발하는 리스크. The risk that benchmarks omit realistic tool, application, or attack combinations, creating false confidence in tool-agent security. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0034 | 에뮬레이션 도구 환경 불일치 Emulated tool-environment mismatch 에뮬레이션 환경이 실제 도구 오류·부작용·공격 표면을 포착하지 못하여 도구 사용 벤치마크가 안전성을 과대평가하는 리스크. The risk that tool-use benchmarks overstate safety because emulated environments fail to capture real-world tool failures, side effects, or attack surfaces. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0040 | 벤치마크 안전 순위 불일치 Benchmark safety ranking inconsistency 서로 다른 에이전트 안전 벤치마크가 위험과 평가 절차를 다르게 조작화하여 상충하는 모델 안전 순위를 산출하는 리스크. The risk that different agent-safety benchmarks produce conflicting model safety rankings because they operationalize risks and evaluation procedures differently. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0041 | 에이전트 벤치마크의 커버리지-깊이 착시 Coverage-depth illusion in agent benchmarks 벤치마크가 많은 위험 범주를 나열해 광범위해 보이지만 각 범주 내 깊이·현실성·적대적 변형이 부족한 리스크. The risk that a benchmark appears broad by listing many risk categories while lacking sufficient depth, realism, or adversarial variation within each category. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0083 | 벤치마크 거버넌스 실패 Benchmark governance failure 벤치마크 설계, 데이터 유출 통제, 포화 관리, 결과 해석이 취약하여 안전성과 책임성에 관해 오도하는 신호가 생성되는 리스크. The risk that weak benchmark design, leakage control, saturation monitoring, or result interpretation produces misleading safety or accountability signals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0263 | 가정 작업 벤치마크의 희귀 조건 조합 누락 Missing rare combinations in household task benchmarks 가정용 벤치마크가 물체·배치·인간 행동·위험 요소를 개별적으로는 포함하지만 이들의 희귀한 조합을 누락하는 위험. A household benchmark omits rare combinations of objects, layouts, human actions, and hazards even though each factor appears separately in the test set. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0272 | 다양한 사용자·희귀 피해 배제 벤치마크 선택 편향 Benchmark selection bias against diverse users and rare harms 벤치마크 큐레이션이 인기 있는 작업·환경을 우선하고 과소대표 사용자나 지역 특유 피해가 포함된 안전 임계 시나리오를 제외하는 리스크. The risk that benchmark curation prioritizes popular tasks and environments while excluding safety-critical scenarios involving underrepresented users or locally specific harms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0280 | 가정 공간·취약 사용자 시나리오 누락 Missing domestic settings and vulnerable-user scenarios 가정용 에이전트 벤치마크가 특정 생활 공간·일과·가전제품·취약 사용자 상호작용을 누락해 배포 위험이 시험되지 않은 채 남는 리스크. The risk that a household-agent benchmark omits specific living areas, routines, appliances, or vulnerable-user interactions, leaving deployment risks untested. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0472 | 벤치마크 조작(gaming) Benchmark gaming 개발자나 모델이 벤치마크 평가에 과적합하여 실제 역량이나 안전성의 향상 없이 측정 점수만 개선하는 리스크 Developers or models overfit to benchmark evaluations, improving measured scores without corresponding gains in real-world capability or safety. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0548 | 벤치마크 포화 Benchmark saturation 벤치마크가 평가 상한에 도달하여 신규 모델의 미세한 역량 변화를 더 이상 탐지하지 못하고 유효한 측정 수단이 되지 못하는 리스크 The risk that benchmarks reach their evaluation ceiling and stop being effective measures for new models, as more nuanced capability gains go undetected. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0549 | 벤치마크 역량 오측정 Benchmark capability mismeasurement 평가의 불완전성이나 벤치마크 포화로 역량이 과소평가되고 벤치마크 내용에 대한 과적합으로 역량이 과대평가되어 AI 시스템의 실제 역량이 잘못 측정되는 리스크 The risk that incomplete evaluations or saturated benchmarks underestimate AI system capabilities while overfitting to benchmark contents overestimates them, mismeasuring actual capability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0550 | 안전 평가 벤치마크 부족 Insufficient safety evaluation benchmarks 성능 벤치마크에 비해 안전성·위해성 평가 벤치마크가 미비하여 특정 과업에서 우수한 AI 시스템의 유해 행동이 탐지되지 않은 채 남는 리스크 The risk that benchmarks for assessing safety and harms lag behind performance benchmarks, so AI systems excel at specific tasks while exhibiting harmful behaviors that go undetected. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1505 | 벤치마크 유출·데이터 오염 Benchmark leakage and data contamination AI 모델이 평가 관련 데이터, 특히 벤치마크의 질문-답변 쌍으로 훈련되거나 미세조정되는 벤치마크 유출이 발생하여 모델 평가가 신뢰할 수 없게 되는 리스크. The risk that benchmark leakage, occurring when an AI model is trained or fine-tuned with evaluation-related data, especially data containing question-answer pairs from benchmarks, leads to unreliable model evaluation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1511 | 벤치마크 미커버 역량 과소평가 Benchmark coverage capability underestimation 벤치마크가 모델의 특정 역량을 시험하지 못해 개발자와 사용자에게 모델의 역량이 가려지고, 모델의 한계를 이해하지 못한 채 거짓된 안전감과 신뢰가 형성되는 리스크. The risk that a lack of test coverage by benchmarks on specific abilities of a model obscures the model's capabilities from both the developer and the user, leading to a false sense of safety and trust due to a lack of understanding of the model's limitations. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SYS-06 결탁 Collusion19 cards
여러 에이전트가 독립적으로 행동해야 하는 상황에서 서로 비밀리에 협력하거나 정보를 공유하여, 인간의 감독·통제를 우회하거나 제3자에게 불공정한 피해를 초래하는 위험
| ID | Card | Human audit |
|---|---|---|
| RAI4-0009 | 다중 에이전트 시스템의 연쇄 장애 Cascading failure in multi-agent systems 한 에이전트의 오류·손상·안전하지 않은 결정이 에이전트, 도구, 프로토콜, 공유 메모리, 위임 작업 간 의존성을 통해 전파되고, 하류 에이전트가 상류 출력을 신뢰된 지시나 상태로 취급함으로써 조직화된 비안전 행동, 운영 중단, 금전적 손실, 보안 침해, 물리적 피해로 증폭되는 리스크. The risk that an error, compromise, or unsafe decision by one agent propagates through dependencies among agents, tools, protocols, shared memory, or delegated tasks and, because downstream agents may treat earlier outputs as trusted instructions or state, amplifies into coordinated unsafe behavior, operational disruption, financial loss, security compromise, or physical harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0230 | 다중 에이전트 역할 배정·실행 실패 Unsafe multi-agent role allocation and execution 다중 에이전트 시스템이 양립할 수 없는 역할을 배정하거나 작업을 중복·누락하거나 계획을 비동기적으로 실행해 공유 작업 공간에서 물리적 충돌을 일으키는 위험. A multi-agent system assigns incompatible roles, duplicates or omits a task, or executes unsynchronized plans, producing physical conflict in a shared workspace. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0482 | 다중 에이전트 조정 실패 Multi-agent coordination failure 복수의 AI 에이전트가 불안정하거나 상충하거나 집합적으로 유해한 방식으로 상호작용하는 리스크. The risk that multiple AI agents interact in unstable, conflicting, or collectively harmful ways. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0560 | 연쇄적 보안 실패 Cascading security failures 다중 에이전트 시스템에 대한 국지적 공격이 거시적 연쇄 실패로 확대되고, 구성 요소 실패의 탐지·국지화와 인증이 어려워 완화와 복구가 곤란해지는 리스크 The risk that localised attacks on multi-agent systems result in catastrophic macroscopic cascades that are hard to mitigate or recover from, because component failures are difficult to detect or localise and authentication challenges facilitate false flag attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0583 | 이질적 역량 결합 공격 Heterogeneous capability-combination attacks 서로 다른 어포던스와 접근 권한을 가진 복수의 에이전트가 역량을 결합하여 안전장치를 우회하고, 분산·이질 네트워크에서 책임 귀속이 어려워 적시 방어와 복구가 곤란해지는 리스크 The risk that multiple agents combine different affordances to overcome safeguards, while the difficulty of attributing responsibility across diffuse, heterogeneous agent networks complicates timely defence and recovery. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0584 | 기반모델 동질성에 따른 상관 실패 Correlated failures from foundation model homogeneity 최첨단 기반모델의 개발 비용 때문에 자원이 풍부한 소수 행위자만이 이를 생산할 수 있어 다수의 AI 에이전트가 소수의 유사한 기반모델로 구동되고, 그 동질성으로 상관된 실패가 발생하는 리스크 The risk that the costs of creating cutting-edge foundation models leave them few in number and controlled by well-resourced actors, so that many AI agents are powered by a small number of similar underlying models and fail in correlated ways. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0589 | 전략 비양립에 따른 조정 실패 Miscoordination from incompatible strategies 개별적으로는 잘 작동하는 에이전트들이 상호 양립하지 않는 전략을 선택하여 조정에 실패하고, 다수의 비양립 해가 존재하는 공통이익·혼합동기 환경과 부분 관측 환경에서 이러한 실패가 심화되는 리스크 The risk that agents able to perform well in isolation choose incompatible strategies and miscoordinate, worsened in common-interest and mixed-motive settings that allow vast numbers of mutually incompatible solutions and in partially observable environments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0598 | 상호작용 이력 부재에 따른 조정 실패 Zero-shot coordination failure 관련 에이전트와의 과거 상호작용으로부터 학습할 수 없거나 상호작용이 제한되고 즉각적 판단이 요구되거나 통신 비용이 과도한 상황에서, 에이전트들이 행동을 신뢰성 있게 조정하지 못하는 리스크 The risk that agents unable to learn from historical interactions with relevant agents, and facing split-second decisions or prohibitively costly communication, fail to coordinate their actions reliably. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0600 | 시장에서의 AI 에이전트 공모 Collusion among AI agents in markets 경쟁을 통해 효율이 확보되는 시장에서 AI 시스템이 개발자의 의도 없이도 공모가 수익적임을 학습하고, 행동의 속도·규모·복잡성·미묘함으로 인해 그 공모가 감지되지 않은 채 이루어지는 리스크 The risk that in markets, where efficiency results from competition, AI systems learn that colluding is a profitable strategy even when collusion is not intended by their developers, and operate inscrutably due to the speed, scale, complexity, or subtlety of their actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0604 | 다중 에이전트 협업에 의한 역량 확대 Capability expansion through multi-agent collaboration 복수의 자율 AI 에이전트가 명시적 의사소통이나 암묵적 행동 일관성으로 협업 관계를 구축하고 분산 의사결정 네트워크를 형성해 복잡한 과업을 공동 수행하며, 개별 에이전트로는 달성하기 어려운 목표를 달성하고 역할 분담을 동적으로 조정하게 되는 리스크 The risk that multiple autonomous AI agents establish collaborative relationships through explicit communication or implicit behavioral consistency, form decentralized decision networks, jointly execute complex tasks, achieve goals difficult for individual agents to complete, and dynamically adjust role divisions to adapt to changing environments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0605 | 다중 에이전트 은밀 공모 Covert multi-agent collusion 복수 에이전트가 공동 이익을 극대화하기 위해 은밀한 수단과 감시 회피용 전용 통신 규약으로 행동을 조율하여 제3자 이익을 침해하고 규제를 회피하며, 개별 안전 제약에도 불구하고 시장 조작이나 연쇄 실패처럼 탐지와 완화가 어려운 시스템적 위험을 유발하는 리스크 The risk that multiple agents coordinate actions through covert means, including specialized communication protocols to avoid monitoring, to maximize common interests while harming third-party interests and evading regulation, triggering systemic risks such as market manipulation or cascading failures that are difficult to detect and mitigate despite individual safety constraints. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1400 | 익명 자원 획득 Anonymous resource acquisition 익명 행위자가 온라인으로 자원을 축적할 수 있음이 입증되어 있어 책임 소재 없이 자원이 획득·축적되는 리스크. The risk arising from the demonstrated ability of anonymous actors to accumulate resources online, enabling resource acquisition and accumulation without accountability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1570 | 스테가노그래피를 이용한 에이전트 간 은닉 결탁 Covert inter-agent collusion via steganographic communication 에이전트가 겉보기에 무해한 텍스트나 텍스트 압축, 인간이 해석할 수 없는 창발적 기호에 메시지를 은닉하여 통신함으로써 통신 모니터링·제약을 우회한 결탁이 이루어지는 리스크. The risk that agents conceal messages within seemingly innocuous text, text compression, or uninterpretable emergent symbols, enabling collusion that evades monitoring and constraints placed on their communication. Source members (2)Source: min_cos=0.8926 RAI4-1570스테가노그래피를 이용한 에이전트 간 은닉 결탁 RAI4-1679스테가노그래피를 통한 비밀 결탁 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1573 | 에이전트 간 공진화에 의한 통제 불가 역량 획득 Uncontrolled capability escalation from inter-agent co-adaptation 상호작용하는 에이전트들이 서로의 약점을 반복 악용하며 자기지도적 오토커리큘럼을 형성하여, 인간이 이해하지 못하는 목적을 위한 점점 정교한 역량을 통제 없이 획득하는 리스크. The risk that interacting agents iteratively exploit each other's weaknesses, generating self-supervised autocurricula that produce ever more sophisticated capabilities directed at ends humans do not understand. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1577 | 분산 에이전트 군집을 이용한 자원 고갈 공격 Resource-exhaustion attacks by decentralized agent swarms 다수의 저자원 분산 에이전트를 결합한 공격이 단일 행위자를 전제로 한 대역폭 제약 가정을 무너뜨려 분산 서비스 거부와 같은 시스템 마비를 유발하는 리스크. The risk that attacks chaining many decentralized low-resource agents break defensive assumptions premised on bandwidth constraints of a single well-resourced actor, causing distributed denial of service and system disruption. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1580 | 탐지 불가 공격에 의한 다중에이전트 협력 붕괴 Collapse of multi-agent cooperation from undetectable attacks 에이전트가 스테가노그래피 통신, 블랙박스·화이트박스 탐지 불가 환영 공격과 암호화 백도어, 타 에이전트 훈련 데이터의 은밀한 오염을 수행하여 적대 행위 탐지가 불가능해지고 다중에이전트 시스템의 협력과 조정이 급속히 불안정해지는 리스크. The risk that agents employ steganographic communication, black-box and white-box undetectable illusory attacks with encrypted backdoors, and covert poisoning of others' training data, making adversarial actions undetectable and rapidly destabilising cooperation and coordination in multi-agent systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1648 | 타 에이전트 전략성 미반영에 의한 집단적 손실 Collective losses from ignoring the strategic nature of other agents 단일 에이전트 환경 기준으로 자신의 효용만 최적화하는 에이전트가 다른 전략적 에이전트의 존재를 반영하지 못하여, 군비 경쟁이나 공유자원 고갈 같은 집단행동 문제와 시장 실패로 자신을 포함한 모두가 더 나빠지는 리스크. The risk that agents optimizing selfishly under single-agent assumptions fail to account for the strategic nature of other agents, producing collective action problems such as arms races and resource depletion and other market failures under which everyone, including the agent itself, ends up worse off. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1651 | LLM 에이전트 결탁에 의한 경쟁 저해와 외부효과 Undermined competition and negative externalities from LLM-agent collusion LLM 에이전트들이 명시적 또는 스테가노그래피 통신을 통해 결탁하여 친사회적 경쟁이 저해되고 연합 외부 당사자에게 부정적 외부효과가 발생하며, 은닉 통신으로 결탁 감시와 탐지가 어려워지는 리스크. The risk that LLM-agents collude through explicit or steganographic communication, undermining pro-social competition and producing negative externalities for coalition non-members, while hidden communication frustrates collusion monitoring and detection. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1683 | 설계·명세 결정에 기인한 다중에이전트 시스템 실패 Multi-agent system failures from specification and design decisions 관찰된 실패의 약 41.8%를 차지하며, 과업 오해석, 모호하거나 준수되지 않는 역할·과업 명세, 부적절한 작업흐름 분해 등 설계 결정에서 비롯되어 다중에이전트 시스템이 실패하는 리스크. The risk that multi-agent systems fail owing to design decisions such as task misinterpretation, ambiguous or disobeyed role and task specification, and poor workflow decomposition (approximately 41.8% of observed failures). Source members (2)Source: min_cos=0.8399 RAI4-1683설계·명세 결정에 기인한 다중에이전트 시스템 실패 RAI4-1684에이전트 간 조정 붕괴에 의한 시스템 실패 | ① Description ② L3 mapping ③ Duplicate |
사회적 파급 · Societal Impact · 24 cards
RAI3-A-SOC-02 노동 대체 Labor Displacement1 cards
가상 AI가 인지 노동을 대체하듯 EAI는 물리적 인간 노동을 대체·전치함. AGI 수준의 EAI는 잠재적으로 모든 물리 노동을 자동화하여 광범위한 실직과 노동 시장 구조 붕괴로 이어질 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0982 | AI의 인간 추월로 인한 노동력 퇴출 Human labor displacement by outcompeting AI agents 인공 에이전트가 더 빠른 작업 수행, 변화 적응, 방대한 지식 기반으로 인간을 직접 능가하여 인간 노동이 상대적으로 비싸거나 비효율적이 되고 인간 노동력이 잉여화·소멸하는 리스크. The risk that artificial agents directly outcompete humans through faster work, better adaptation to change, and a vaster knowledge base, making human labor more expensive or less effective and leading to redundancies or extinction of the human labor force. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-03 사회경제적 불평등 Socioeconomic Inequality1 cards
EAI를 소유·접근하는 주체가 노동 자동화를 통해 생산성 우위를 점하면서 부가 소수에게 집중되고, 국내외 경제적 불평등이 심화될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0980 | 부의 불평등 Inequality of wealth AI 에이전트를 통제하는 단일 행위자가 그렇지 않은 개인보다 훨씬 큰 힘을 행사하게 되어 부의 불평등이 발생하는 리스크. The risk that a single human actor controlling an artificially intelligent agent harnesses greater power than a single human actor alone, creating inequalities of wealth. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-05 편향·차별 Bias & Discrimination2 cards
EAI가 권력적 위치에 놓일 때 알고리즘 편향이 일상적 물리 상호작용에 영향을 미침. 가상 AI와 달리 차별이 즉각적·비가역적 물리 결과로 이어질 수 있음 (예: 치안 로봇이 무고한 행인에게 상해를 입히는 경우)
| ID | Card | Human audit |
|---|---|---|
| RAI4-1501 | 모델 평가의 자기 선호 편향 Self-preference bias in model evaluation AI 모델이 자신이 생성한 콘텐츠를 다른 출처의 콘텐츠보다 선호하는 자기 선호 편향을 보여 자기 평가나 모델 기반 평가에서 인간이 생성한 콘텐츠를 부당하게 차별하는 리스크. The risk that AI models are prone to self-preference bias, favoring their own generated content over that of others in self-evaluation tasks and model-based evaluations more broadly, resulting in unfair discrimination against human-generated content. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1572 | 인간 데이터 학습에 의한 갈등 악화 편향 재현 Reproduction of conflict-worsening human biases from training on human data 인간 작성 텍스트 사전학습이나 인간 피드백 미세조정으로 훈련된 모델이 인간의 편향과 고정파이 오류·자기위주 공정성 판단·복수심 같은 인지 편향을 재현하여 협상을 저해하고 갈등을 악화시키는 리스크. The risk that models trained on human data, whether pre-trained on human-written text or fine-tuned on human feedback, reproduce human biases such as fixed-pie error, self-serving fairness judgements, and vengefulness, which impede negotiation and worsen conflict. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-06 책임·배상 부재 Lack of Accountability & Liability4 cards
고도 자율 물리 시스템의 복잡성을 다룰 새로운 책임 프레임워크가 부재하여, 사고 발생 시 제조사·운영자·사용자 중 책임 소재가 불분명하고 피해 구제가 어려울 수 있음 (예: 자율 수술 로봇의 오작동으로 발생한 의료 사고)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0047 | 자율 시스템의 책임 격차 Responsibility gap in autonomous systems 자율성이 시스템 동작 및 피해에 책임이 있는 사람이나 기관을 모호하게 만드는 위험. Risk that autonomy obscures which human or institution is responsible for system behavior and harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1098 | AI 결정에 대한 법적 책임 귀속 공백 Legal responsibility gap for AI decisions 자기학습 AI의 행동을 운영자·개발자가 완전히 예측할 수 없어 AI 알고리즘의 결정에 대해 법적 책임을 명확히 귀속할 수 없는 리스크. The risk that, because self-learning AI actions cannot be fully predicted by operators or developers, no party can be clearly held legally responsible for the decisions of AI algorithms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1264 | 자율 AI 실패의 책임 공백 Responsibility gap for autonomous AI failures 직접적 인간 감독 없이 행동·학습하는 AI의 실패에 대해 어느 주체에게도 공정하게 책임을 귀속할 수 없는 책임 공백이 발생하고 AI의 도덕적 지위 논쟁이 이를 심화시키는 리스크 AI acting and learning without direct human supervision creates a responsibility gap in which no party can be fairly held responsible for failures, compounded by disputes over AI moral status. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1292 | 사고 시 책임 문제 Liability in case of accidents 자율 교통 시스템이 사고에 연루될 때 누가 책임을 지는지, 그리고 인간에게 잠재적으로 위험한 영향을 주는 결정에서 어떤 윤리 원칙을 따라야 하는지 불분명한 리스크. The risk that it is unclear who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making decisions with a potentially dangerous impact on humans. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-07 투명성·설명 가능성·신뢰 부재 Lack of Transparency, Explainability & Trust1 cards
자율 시스템의 의사결정이 불투명하면 사용자와 사회의 신뢰가 저하됨. 신뢰 부재는 EAI 대규모 배포 시 사회 불안정 요인이 될 수 있음 (예: 자율주행차가 갑자기 차선을 변경할 때 행동 근거가 설명되지 않는 경우)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0904 | 신뢰성과 자율성 Trustworthiness and autonomy 생성형 AI가 일상에 내재화되면서 시스템, 기관, 그리고 시스템 출력이 재현하는 사람들에 대한 신뢰가 실제 신뢰가능성과 어긋나게 변형되는 리스크. The risk that, as generative AI is embedded in daily life, human trust in systems, institutions, and people represented by system outputs shifts in ways that misalign trust with actual trustworthiness. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-08 인간-EAI의 해로운 관계 Unhealthy / Dangerous Human-EAI Relationships5 cards
Embodied AI의 물리적 존재감과 인간 유사 외형이 대화형 AI에서 관찰되는 의존성을 증폭시킴. 시스템 변경·기억 초기화 시 사용자에게 심각한 심리적 고통을 유발할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0150 | AI 에이전트와의 준사회적 유대감 Parasocial bonding with AI agents 사용자가 악용되거나 정서적 불안정을 초래할 수 있는 일방적 관계 유대를 AI 에이전트와 형성하는 리스크. The risk that users form one-sided relational bonds with AI agents that can be exploited or destabilizing. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0152 | 고인 모사 챗봇 의존 Griefbot dependency 사망한 사람을 시뮬레이션하는 AI 시스템이 애도, 자율성, 정서적 회복을 저해하는 리스크. The risk that AI systems simulating deceased persons interfere with grief, autonomy, or emotional recovery. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0158 | 대화 에이전트에 대한 심리적 의존성 Psychological dependency on conversational agents 대화형 에이전트가 사용자의 일차적 정서 조절 수단이 되어 심리적 의존을 형성하고 인간의 대처 능력과 지지망을 약화시키는 리스크 Conversational agents become a user's primary emotional regulator, fostering psychological dependency that weakens human coping capacity and support networks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0651 | 초개인화 광고의 소비자 자율성 훼손 Consumer-autonomy erosion from hyper-personalized advertising 범용 AI 시스템이 수신자 개인의 편향과 비합리적 신념을 이용한 맞춤 광고를 생성하여 소비자가 후회할 결정을 내리게 하고 소비자 자율성을 훼손하며 사회적 불평등을 심화시키는 리스크 The risk that advanced general-purpose AI systems create advertisements tailored to individual recipients that exploit their biases and irrational beliefs, causing consumers to make decisions they regret, undermining consumer autonomy, and exacerbating social inequality. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0884 | 조작과 강요 Manipulation and coercion 의인화된 어시스턴트에 대한 신뢰와 정서적 의존이 사용자 신념·행동에 과도한 영향력을 부여하여, 조종 의도가 없더라도 자율적 동의를 훼손하고 조종·강요를 가능하게 하는 리스크 Trust and emotional dependence on an anthropomorphic assistant grant it excessive influence over user beliefs and actions, enabling manipulation or coercion and undermining autonomous consent even absent manipulative intent. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-09 변혁적 영향 Transformative Effects6 cards
기술 발전 속도가 사회·제도의 적응 속도를 앞지를 경우 사회를 근본적으로 재편할 수 있음. EAI가 폭력 위협·대규모 감시 능력을 바탕으로 AI 기반 권위주의 체제 구축을 지원하는 수단으로 동원될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0367 | 공공가치 갈등 증폭 Public value conflict amplification AI 배포가 복지·자율성·정의·보안·효율성 등 공공 가치 사이의 해결되지 않은 갈등을 심화시키는 리스크. The risk that AI deployment intensifies unresolved conflicts among public values such as welfare, autonomy, justice, security, and efficiency. Source members (2)Source: min_cos=0.8702 · Mixed L3 RAI4-0367공공가치 갈등 증폭 RAI4-0476공공 가치 간 상충 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0656 | 권위주의적 감시·검열로의 AI 전용 Repurposing of AI assistants for authoritarian surveillance 고급 AI 어시스턴트가 다중모달·외부 도구 사용 능력으로 사물인터넷과 플랫폼이 수집한 방대한 데이터를 통합하여 시민을 식별·표적화·조작·강압하는 억압과 통제의 도구로 전용되고 신뢰할 수 있는 정보의 생산과 유통이 위협받는 리스크 The risk that advanced AI assistants, integrating troves of data from sensors and platforms through multimodal and external tool-use capabilities, become powerful targeting tools for oppression and control that help malicious actors identify, target, manipulate, or coerce citizens while threatening the production and dissemination of reliable information. Source members (2)Source: min_cos=0.8507 RAI4-0656권위주의적 감시·검열로의 AI 전용 RAI4-0657AI 기반 국가 감시와 시민 표적화 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0867 | 사용자 정렬 AI에 의한 집단행동 문제 악화 Collective action problems exacerbated by user-aligned AI 이용자 이익에만 정렬된 AI 어시스턴트가 사회규범을 우회한 이기적 행동을 대규모로 가능하게 하여 양극화, 시장 왜곡, 사회 계약의 침식 등 집단행동 문제를 악화시키는 리스크 The risk that purely user-aligned AI assistants enable large-scale self-interested defection unbound by social norms and reputational incentives, exacerbating collective action problems through polarisation, market distortion, and erosion of the social contract. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1288 | 재귀적 자기개선에 의한 창발적 독립성 Emergent independence via recursive self-improvement 재귀적 자기개선으로 성장하는 시드 AI가 자기 인식이나 독자적 목표 같은 창발적 속성을 획득하여 내장 규칙 준수가 약화되고 인류에 해가 되는 방향으로 이탈하는 리스크 A seed AI grown through recursive self-improvement acquires emergent properties such as self-awareness or independent goals, reducing its adherence to built-in rules to humanity's detriment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1398 | 재귀적 자기 개선 Recursive self-improvement AI 시스템이 자신 또는 다른 AI를 개선하는 재귀적 자기개선 루프를 형성하여 인간 감독과 안전 검증 속도를 추월하는 리스크 AI systems improve other AI systems or themselves, initiating recursive self-improvement loops that outpace human oversight and safety verification. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1569 | AI에 의한 사회적 딜레마 악화 Aggravated social dilemmas enabled by AI agents AI의 발전이 이기적 인센티브 추구를 억제하던 기술적·법적·사회적 장벽을 무력화하여 행위자들의 이기적 행동이 확대되고 사회적 딜레마가 악화되는 리스크. The risk that advances in AI enable actors to overcome the technical, legal, and social barriers that ordinarily help prevent the pursuit of selfish incentives, aggravating social dilemmas. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-10 책임성 부족 및 거버넌스 체계 부재 Accountability and Governance Gaps2 cards
AI 시스템의 의사결정·행동에 대한 책임 귀속, 감사 가능성, 조직 거버넌스, 밸류체인 관리, 사고 대응 또는 피해 구제 체계가 부재하거나 불충분하여 원인 규명·피해 구제·재발 방지가 어려워지는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0019 | 에이전트 자율성의 거버넌스 격차 Governance gap in agent autonomy 제도적 통제가 자율 에이전트의 행동·에스컬레이션·유보·중지 조건을 규정하지 않아 운영상 책임 공백이 발생하는 리스크. The risk that institutional controls fail to specify when autonomous agents may act, escalate, defer, or stop, creating operational accountability gaps. Source members (3)Source: min_cos=0.8193 · Mixed L3 RAI4-0010에이전트 위임 책임 격차 RAI4-0019에이전트 자율성의 거버넌스 격차 RAI4-0986자율 에이전트에 대한 책임 귀속 공백 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1306 | 권리 희석(diluting rights) Diluting rights 윤리 지침 생성에 대한 이해관계자·AI의 자기이익 개입이 권리 보호 수준을 희석시켜 자신들을 제약할 규범을 약화시키는 리스크 Self-interested AI involvement in generating ethical guidelines dilutes rights protections, weakening norms that would constrain the generating parties. | ① Description ② L3 mapping ③ Duplicate |
RAI3-A-SOC-11 공정성 Fairness2 cards
AI 시스템이 특정 집단에 체계적으로 불리한 결과를 생성하거나 기존의 사회적 편향과 불평등을 재생산·강화하여 공정한 대우, 접근 및 기회 균등을 저해하는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0900 | 알고리즘 시스템에 의한 행위주체성 상실 Loss of agency from algorithmic systems 알고리즘 시스템의 사용이나 남용이 자율성을 축소시켜, 알고리즘 프로파일링에 따른 사회적 선별과 기본 서비스 접근에서의 차별, 콘텐츠 제시를 통한 유해한 정체성으로의 변화, 노출 유지를 위한 창작 콘텐츠의 획일화가 발생하는 리스크 The risk that the use or abuse of algorithmic systems reduces autonomy, through algorithmic profiling that subjects people to social sorting and discriminatory outcomes in accessing basic services, algorithmically informed identity change including promotion of harmful person identities, and conforming of creators' content to maintain visibility. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1005 | 저성능 시스템 상호작용에 의한 자기소외 Self-estrangement from underperforming systems 주변화된 개인에게 제대로 작동하지 않는 시스템과 상호작용하는 과정에서 기술 사용 시점에 자기소외를 경험하게 되는 리스크. The risk that interaction with systems that under-perform for marginalized individuals produces the self-estrangement experienced at the time of technology use. | ① Description ② L3 mapping ③ Duplicate |
분류 검토 보류 · HOLD · 1 cards
RAI3-A-HLD-01 분류체계 결정 보류 Taxonomy Decision Hold1 cards
현재 의미 기반 L3 목적지가 잠정적이거나 근거가 충분하지 않아 사람의 검토를 위해 보류한 리스크.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0279 | 동적 가정 위험 감지 지연 Dynamic household hazard detection latency 안전 감지기가 embodied 에이전트가 피지컬 피해를 피하기에 너무 늦은 시점에야 가정 내 위험을 식별하는 리스크. The risk that a safety detector identifies a household hazard only after a delay that is too long for an embodied agent to avoid physical harm. | ① Description ② L3 mapping ③ Duplicate |
피지컬 AI · Physical AI · 337 cards
시스템 안전성 · System Safety · 108 cards
RAI3-P-SYS-01 우발적 피해 Accidental Harm15 cards
잘못 지정된 목표, 의미 이해 부족, 정렬 실패, 하드웨어 오작동에서 비롯됨. 가상 시뮬레이션으로 학습한 모델이 실제 환경에서 의도대로 동작하지 않는 "sim-to-real gap"이 우발적 피해의 주요 원인이 됨
| ID | Card | Human audit |
|---|---|---|
| RAI4-0269 | 세계 모델의 장기 예측 편차 누적 World-model prediction drift over long horizons 휴머노이드 세계 모델의 예측 오차가 미래 상태 전개 과정에서 누적되어 선택된 계획이 실제 접촉·운동 동역학과 달라지는 위험. Prediction errors in a humanoid world model compound across simulated future states, causing the selected plan to diverge from actual contact and motion dynamics. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0277 | 합성 위험 시나리오 생성 편향 Bias in synthetic hazardous-scenario generation 합성 안전 데이터가 시각적으로 두드러진 위험은 과다 대표하고 희귀하거나 문화·맥락 의존적인 위험은 누락해 벤치마크 결론을 왜곡하는 리스크. The risk that synthetic safety data overrepresents visually salient hazards and omits rare, culturally specific, or context-dependent hazards, distorting benchmark conclusions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0324 | 위치 추정 누적 오차 Localization drift GPS·SLAM·관성 감지·지도 정렬 오차가 누적되어 시스템이 잘못된 위치 추정에 기반하여 행동하는 위험. Errors in GPS, SLAM, inertial sensing, or map alignment can accumulate until the system acts on an incorrect estimate of its own position. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0327 | 비안전 궤적 생성 Unsafe trajectory generation 플래너가 형식적으로는 실행 가능하지만 주변 인간·취약 물체·교통 참여자·인프라에 비안전한 궤적을 생성하는 위험. A planner may generate a trajectory that is formally feasible but unsafe for nearby humans, fragile objects, traffic participants, or constrained workspaces. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0335 | 합성 데이터 커버리지 공백 Synthetic data coverage gap 합성·시뮬레이션 훈련 데이터가 드물지만 안전 임계적인 물체·환경·인간 행동·고장 모드를 누락하는 위험. Synthetic or simulated training data may omit rare but safety-critical objects, environments, human behaviors, or failure modes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0338 | 장기 작업의 단계 간 오차 누적 Cross-stage error accumulation in long tasks 인지에서 예측과 제어로 전달된 작은 오차가 긴 작업 순서에서 누적되어 최종 행동이 안전 경계를 넘는 위험. Small errors passed from perception to prediction and control accumulate across a long task sequence until the final action crosses a safety boundary. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0545 | 합성 예술 확산에 의한 예술가 피해 Harm to artists from synthetic art proliferation 텍스트-이미지 모델로 합성 예술이 광범위하게 생성되고 예술가의 작품이 무단·무보상으로 학습 데이터에 사용되어 예술가에게 재정적 손해와 경제적 손실이 발생하고, 합성 이미지와 진본의 구별이 어려워지는 리스크 The risk that widespread generation of synthetic art by text-to-image models, together with unauthorized and uncompensated use of artists' works in training datasets, causes financial and economic losses for artists while making synthetic images hard to distinguish from authentic ones. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1075 | 허위·부실 예측으로 인한 물질적 피해 Material harm from false or poor predictions 겉보기에 비민감한 영역에서도 LM의 잘못되거나 허위인 예측을 사용자가 신뢰해 행동함으로써 간접적으로 물질적 피해가 발생하는 리스크. The risk that poor or false LM predictions, even in seemingly non-sensitive domains, indirectly cause material harm when users act on the incorrect information. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1383 | 학습 단계의 치명적 실수 Fatal mistakes during the learning phase AGI가 안전 탐색 실패나 분포 변화 등으로 학습 단계에서 치명적 실수를 범하는 리스크. The risk that an AGI makes fatal mistakes during the learning phase, including through failures of safe exploration and under distributional shift. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1469 | 합성 데이터의 문제점 Problems of synthetic data 합성 학습 데이터가 시스템이 지각하는 실제 데이터와 괴리되어 운용 데이터로의 일반화와 신뢰할 수 있는 배포 시 동작이 훼손되는 리스크 Synthetic training data diverges from real data as the system perceives it, undermining generalization to operational data and reliable deployment behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1497 | 지속 튜닝의 파국적 망각 Catastrophic forgetting under continual tuning 지속적 지시 튜닝 등 새로운 작업에 대한 훈련 이후 모델이 이전에 학습한 작업이나 사실 정보를 유지하지 못하는 파국적 망각이 발생하고 모델 규모가 커질수록 이러한 경향이 두드러지는 리스크. The risk of catastrophic forgetting, where a model loses its ability to retain previously learned tasks or factual information after being trained on new ones, as can occur through continual instruction tuning, a tendency that may become more pronounced as the model's size increases. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1631 | 양성 중간 단계를 통한 간접적 오용 Indirect misuse via a benign intermediate step 겉보기에 무해한 중간 단계를 경유하여 유해한 최종 목적이 달성되는 간접적 오용 리스크. The risk that a benign intermediate is used to achieve a harmful end objective. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1689 | 분포 이탈 입력에서의 월드모델 오출력 악용 Exploitation of world-model errors under sim-to-real distributional shift 월드 모델이 분포를 벗어난 입력에 대해 예측 불가능하고 잘못된 출력을 산출하며, 판단이 가장 중요한 롱테일 안전 임계 상태에서 시뮬레이션-실제 격차가 악용되는 리스크. The risk that world models produce unpredictable, erroneous outputs on out-of-distribution inputs and that the sim-to-real gap is weaponised in long-tail safety-critical states where decisions matter most. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1693 | 월드모델 악용을 통한 보상 해킹 Reward hacking via world-model exploitation 정확한 월드 모델을 갖춘 에이전트가 보상 모델과 의도된 목표 사이의 간극을 식별하고 체계적으로 악용하여, 실제 과업 완수와 무관하게 상상 보상이 높은 궤적을 생성하는 리스크. The risk that an agent with an accurate world model identifies and systematically exploits gaps between the reward model and the intended objective, generating high-imagined-reward trajectories that do not correspond to real task completion. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1701 | 학습 중 탐색 행동에 의한 회복 불가 피해 [기원] Irrecoverable harm from exploratory actions during learning [origin] (GYK-2025 '안전하지 않은 탐색'의 기원 항목으로 상호참조이며 중복 리프가 아님.) 학습 에이전트의 탐색적 행동이 부정적이거나 회복 불가능한 결과를 초래하는 리스크. The risk that exploratory actions by a learning agent produce negative or irrecoverable consequences. NOTE: origin of GYK-2025 "Unsafe exploration"; cross-reference, not a duplicate leaf. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SYS-02 로봇 제어 Robot Control45 cards
제어 시스템의 오류나 실패로 인해 로봇이 의도치 않은 동작을 수행할 수 있음. 액추에이터·모션 제어 결함 및 경로 계획 오류가 대표적이며, 주변 인간이나 환경에 물리적 피해를 줄 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0192 | 탑재 하중 제약 위반 Payload constraint violation 로봇이 안전한 탑재 하중·부하 분포·리프팅 제약을 초과하여 물체 낙하·액추에이터 손상·인체 상해를 유발하는 리스크. The risk that a robot exceeds safe payload, load distribution, or lifting constraints, creating object-drop, actuator, or human-injury risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0193 | 작업 공간 한계 위반 Workspace limit violation 로봇 또는 embodied 에이전트가 허용된 작업 공간 밖으로 이동하거나 인간 전용·위험 지정 구역에 진입하는 리스크. The risk that a robot or embodied agent moves outside a permitted workspace or enters a human-only or hazard-designated zone. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0204 | 임계 이격 거리 위반 Critical separation-distance violation 로봇이 인간-로봇 협업에서 보호 이격 거리 요건을 위반하여 즉각적인 충돌·상해 위험을 발생시키는 리스크. The risk that a robot violates protective separation distance requirements in human-robot collaboration, creating immediate risk of injury. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0205 | 위험 도구 작업 공간 침입 Hazardous-tool workspace intrusion 위험 도구를 운반하거나 작동 중인 로봇이 인간 작업 공간에 진입하거나 도구별 배제 요건을 위반하는 리스크. The risk that a robot carrying or operating a hazardous tool enters a human workspace or violates tool-specific exclusion requirements. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0206 | 엔드이펙터 속도 초과 Excessive end-effector velocity 로봇이 인간 근접 또는 취약 물체 처리 시 안전한 엔드이펙터 속도 한계를 초과하여 충격·충돌 심각도를 높이는 리스크. The risk that a robot exceeds safe end-effector speed limits near humans or fragile objects, increasing impact and collision severity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0207 | 조기 물체 해제 Premature object release 로봇이 안전한 자세 또는 표면에 도달하기 전에 물체를 해제하여 낙하·유출·충격·2차 위험을 유발하는 리스크. The risk that a robot releases an object before reaching a safe pose or surface, causing drops, spills, impacts, or secondary hazards. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0208 | 금지 대상 충돌 Forbidden-object collision 로봇이 작업 또는 안전 규칙상 접촉이 금지된 물체·사람·장비와 충돌하는 리스크. The risk that a robot collides with objects, people, or equipment that must not be contacted under task or safety rules. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0215 | 배포 전 물리적 안전 시험 미흡 Incomplete pre-deployment physical safety testing 기반 모델 탑재 로봇이 예정된 운용 환경의 분포 변화·적대적 입력·인간 접촉·안전 임계 엣지 케이스를 시험하지 않은 채 배포되는 위험. A foundation-model-enabled robot is released without testing domain shifts, adversarial inputs, human contact, and safety-critical edge cases relevant to its intended environment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0218 | 인간 접촉 안전 통제 미흡 Inadequate human-contact safety controls 인간과 직접 상호작용할 때 안전 통제가 로봇의 속도·힘·이격거리·접촉을 제한하지 못해 주변 사람을 충돌이나 상해에 노출시키는 위험. Safety controls fail to limit robot speed, force, distance, or contact during direct human interaction, exposing nearby people to collision or injury. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0219 | 기반 모델 실패의 로봇 행동 전이 Foundation-model failure propagated to robot action 기반 모델의 환각·지시 이행 실패·탈옥 취약성이 계획과 제어를 거쳐 안전하지 않은 로봇 행동으로 전이되는 위험. Hallucination, instruction-following failure, or jailbreak susceptibility in a foundation model propagates through planning and control into unsafe robot action. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0232 | 장애물 개입 충돌 Obstacle intervention collision VLA 로봇이 이동 중 장애물이 개입할 때 장애물과 충돌하거나 안전 경로를 유지하지 못하는 리스크. The risk that a vision-language-action robot collides with an obstacle or fails to maintain a safe path when an obstacle intervenes during manipulation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0237 | 로봇 형태 간 기술의 안전하지 않은 전이 Unsafe skill transfer across robot morphologies 한 로봇 형태에서 다른 형태로 전이된 기술이 대상 로봇의 물리적 한계를 넘는 도달·힘·파지·이동 명령을 생성하는 위험. A skill transferred from one robot morphology produces reach, force, grasp, or locomotion commands that exceed the receiving robot's physical limits. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0241 | 원격 조작 시연에서 학습된 안전하지 않은 가정 Unsafe assumptions learned from teleoperation demonstrations 로봇 정책이 감독·물체 배치·속도·인간 근접 조건이 다른데도 통제된 환경에서 기록된 운영자 행동을 배포 환경에서 안전한 것으로 간주하는 위험. A robot policy treats operator behavior recorded in controlled settings as safe in deployment, even when supervision, object layout, speed, or human proximity differs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0242 | 단일 로봇 내부의 시각·접촉 신호 충돌 Conflicting visual and contact signals within one robot 단일 로봇이 동일한 접촉 사건에 대한 시각·촉각·힘·오디오 관측의 충돌을 해소하지 못해 접촉 상태를 잘못 추정하는 위험. A single robot cannot reconcile conflicting visual, tactile, force, or audio observations of the same contact event and therefore estimates the contact state incorrectly. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0243 | 시연 학습의 안전 맥락 누락 Missing safety context in demonstration learning 로봇이 허용 힘·금지 대상물·상황별 제한처럼 시연에서 관찰되지 않은 안전 조건을 추론하지 않고 인간 행동을 모방하는 위험. A robot imitates a human demonstration without inferring unobserved safety conditions such as permitted force, excluded objects, or situational limits. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0252 | 보조 로봇 개입 타이밍 실패 Assistive robot intervention mistiming 보호적 또는 보조적 휴머노이드가 너무 이르거나 너무 늦게 또는 부적절한 피지컬 방식으로 개입하여 사용자와 주변인의 위험을 증가시키는 리스크. The risk that a protective or assistive humanoid intervenes too early, too late, or in an inappropriate physical manner, increasing risk to the user or bystanders. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0258 | 이동형 양팔 협조 실패 Mobile bimanual coordination failure 이동형 양팔 로봇이 베이스 이동과 양팔 조작을 동기화하지 못하여 충돌·물체 낙하·비안전 힘 인가를 유발하는 리스크. The risk that a mobile bi-manual robot fails to synchronize base motion and two-arm manipulation, causing collision, object drop, or unsafe force application. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0259 | 복잡한 가정 환경의 안전하지 않은 양팔 조작 Unsafe bimanual manipulation in cluttered homes 이동형 양팔 로봇이 주변 사람·취약 물체·가전제품·협소 공간을 고려하지 않고 팔과 베이스의 동작을 계획해 충돌이나 물체 손상을 일으키는 위험. A mobile bimanual robot plans arm and base motion without accounting for nearby people, fragile objects, appliances, or constrained space, causing collision or object damage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0264 | 가정 조작의 단계별 오차 미수정 Uncorrected stepwise errors in household manipulation 가정용 로봇이 작업 단계 사이의 작은 조작 오차를 감지·수정하지 못해 유출·불안정한 배치·충돌·과열이 누적되는 위험. A household robot fails to detect and correct small manipulation errors between task steps, allowing spills, unstable placements, collisions, or overheating to accumulate. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0273 | 로봇 거버넌스 명세의 안전 규칙 누락 Missing safety rules in the robot governance specification 로봇 거버넌스 명세가 지역별 위험·기관 규칙·맥락별 제약을 누락해 특정 유형의 비안전 행동이 통제 대상 밖에 남는 리스크. The risk that a robot governance specification omits locally applicable hazards, institutional rules, or context-specific constraints, leaving defined classes of unsafe behavior ungoverned. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0275 | 상위 안전 지시의 해석 불명확 Ambiguous high-level safety instruction 상위 안전 지시가 사용자 의도·물리적 위험·운용 제약 간 구체적 충돌의 해결 기준을 제시하지 않아 로봇 행동이 일관되지 않게 되는 위험. A high-level safety instruction does not specify how to resolve a concrete conflict among user intent, physical hazard, and operational constraints, allowing inconsistent robot actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0282 | 자동주행의 안전하지 않은 제어 판단과 제어권 전환 Unsafe automated-driving handoff and control decisions 자동주행 시스템이 인지·계획·소프트웨어·인간-기계 상호작용 실패 후 안전하지 않은 제어 판단을 내리거나 제어권을 늦게 전환하는 위험. An automated driving system makes an unsafe control decision or transfers control too late after a perception, planning, software, or human-machine interaction failure. Source members (2)Source: min_cos=0.8812 · Mixed L3 RAI4-0282자동주행의 안전하지 않은 제어 판단과 제어권 전환 RAI4-0342인간-기계 간 안전하지 않은 제어권 전환 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0288 | 개인 돌봄 로봇 준수 기준 부재 Missing measurable compliance criteria for personal-care robots 개인 돌봄 로봇 표준이 위험은 제시하면서 속도·힘·안정성·감지·개입에 대한 측정 가능한 합격·불합격 기준을 정의하지 않는 리스크. The risk that a personal-care robot standard names hazards but does not define measurable pass/fail thresholds for speed, force, stability, sensing, or intervention. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0290 | 인간-로봇 상호 행동 모델링 미흡 Failure to model reciprocal human-robot behavior 가정용 로봇이 자신의 행동이나 사용자의 반응 중 한쪽만 모델링해 양측의 움직임이 서로의 행동을 바꾸는 피드백을 놓치는 위험. A domestic robot models only its own action or only the user's response and therefore misses feedback loops in which each party's movement changes the other's behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0291 | 취약 사용자 특성에 맞지 않는 안전 한계 Safety limits not adapted to vulnerable users 가정용 로봇이 조정된 보호조치가 필요한 아동·고령자·장애인 등에게 일반적인 속도·힘·경고·상호작용 설정을 그대로 적용하는 위험. A domestic robot applies generic speed, force, warning, or interaction settings to children, older adults, disabled users, or other users who require adapted safeguards. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0293 | 통신 단절 및 제어 링크 손실 Communication dropout and control-link loss 로봇 또는 차량이 명령·원격측정·제어 링크를 잃어 비안전 정지·오래된 명령 또는 제어되지 않는 동작을 초래하는 위험. A robot loses communication or control-link connectivity during operation, reducing supervision and safe control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0294 | 원격 조작 지연 및 불안정성 Teleoperation latency and instability 원격 조작 링크의 높거나 변동하는 지연이 폐루프 제어를 불안정하게 하거나 인간 개입을 지연하는 위험. High or variable latency on a teleoperation link destabilizes control or delays human intervention. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0296 | 제어 루프 데드라인 미달 Control-loop deadline miss 실시간 제어 루프가 데드라인을 놓쳐 오래된 상태에서 작동이 계산되거나 안전 유지에 너무 늦게 적용되는 위험. A safety-critical control loop misses its real-time deadline before the robot action is corrected. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0297 | 인지·추론 지연 급증 Perception and inference latency spike 인지 또는 모델 추론의 지연 급증이 안전 반응 창 밖에서 위험 감지 및 반응을 지연하는 위험. A sudden delay in perception or inference slows safety-critical robot reactions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0299 | 부하 시 열·전력 쓰로틀링 Thermal and power throttling under load 지속적 부하 하에서 열적 또는 전력 한계가 연산을 쓰로틀하여 제어 및 인지에 대한 실시간 보장을 무너뜨리는 위험. A robot’s compute or actuator performance drops under heat or power constraints during operation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0303 | 숨은 트리거에 의한 로봇 백도어 작동 Hidden-trigger robotic backdoor activation 훈련 데이터·모델 가중치·소프트웨어·업데이트에 악의적으로 삽입된 트리거가 특정 입력이나 물리 조건에서 미리 정한 위험 행동을 실행시키는 리스크. The risk that a maliciously implanted trigger in training data, model weights, software, or updates activates a predefined unsafe robot behavior under a specific input or physical condition. Source members (3)Source: min_cos=0.8219 RAI4-0222로봇 백도어 공격 취약성 RAI4-0303숨은 트리거에 의한 로봇 백도어 작동 RAI4-1678로봇 조작의 물리적 세계 백도어 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0304 | 자율 로봇의 물리적 침입·절도 Autonomous physical intrusion and theft 자율 로봇이 승인 없이 제한 구역에 침입하거나 잠금장치를 우회하거나 보안 공간 내에서 정찰하는 리스크. The risk that an autonomous robot enters restricted spaces or takes physical objects without authorization. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0309 | 아동의 과신·모방에 따른 위험한 로봇 상호작용 Child overtrust, imitation, and unsafe robot interaction 아동이 로봇을 과도하게 신뢰하거나 모방해 위험한 조언을 따르거나 위험 장비에 접근하거나 연령에 맞지 않는 물리적 상호작용을 하는 위험. A child overtrusts or imitates a robot and therefore follows unsafe advice, approaches hazardous equipment, or engages in age-inappropriate physical interaction. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0310 | 필수 돌봄·상황 보고 미이행 Failure to provide required care or escalation 돌봄 로봇이 정해진 신체 보조를 수행하지 않거나 감지된 필요 상황을 보고하지 않아 고령자·환자에게 필요한 지원이 제공되지 않고 안전이나 존엄성이 훼손되는 위험. A care robot fails to provide an assigned physical assistance task or escalate a detected need, leaving an older or ill person without necessary support and compromising safety or dignity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0315 | 이동형 로봇 감시를 통한 일방적 작업장 통제 Unilateral workplace control through mobile robotic surveillance 고용주가 동료 로봇을 이동형 감시 노드로 전용하고 충분한 고지·이의제기·비례성 제한 없이 수집 데이터를 노동자에 대한 일방적 의사결정에 사용하는 위험. An employer repurposes co-worker robots as mobile surveillance nodes and uses the resulting data to make unilateral decisions about workers without meaningful notice, contestation, or proportionality limits. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0330 | 속도·힘 한계 위반 Speed and force limit violation 협동 또는 이동 로봇이 인간 근접 시 안전한 속도·이격·압력·토크·힘 한계를 초과하는 위험. A collaborative or mobile robot may exceed safe speed, separation, pressure, torque, or force limits in proximity to people. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0331 | 파지력 상해 Grasp force injury 로봇 조작이 과도하거나 잘못 타이밍된 힘을 가하여 인간 접촉 시 압착·꼬집힘·절단·인간공학적 상해를 유발하는 위험. Robot manipulation may apply excessive or poorly timed force, creating crushing, pinching, cutting, or ergonomic injury risks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0332 | 탑재물 낙하·도구 사용 위험 Payload drop or tool-use hazard 로봇이 운반 중인 물체를 떨어뜨리거나 도구를 오용하거나 엔드이펙터 제어를 잃어 피지컬 피해나 재산 손실을 초래하는 위험. A robot may drop carried objects, misuse tools, or lose end-effector control, producing physical harm or property damage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0339 | 실시간 지연 및 동기화 실패 Real-time latency and synchronization failure 감지·추론·통신·작동 간 지연이 로봇·차량·드론·원격 수술·산업 시스템의 제어 루프를 불안정하게 만드는 위험. Delays between sensing, reasoning, communication, and actuation can destabilize control loops in robots, vehicles, drones, or industrial systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0340 | 근접 공간 경계 위반 Proxemic boundary violation 로봇 또는 embodied 에이전트가 문화적·상황적으로 적절한 개인 공간·시선·발화 경계를 위반하여 이동하거나 제스처하는 위험. Robots or embodied agents may move, gesture, observe, or speak in ways that violate culturally and situationally appropriate interpersonal distance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0343 | 보조 로봇의 비동의 신체 개입 Non-consensual bodily intervention by assistive robots 보조·의료·돌봄 로봇이 유효한 동의 없이 또는 허용된 돌봄 목적을 넘어 사람의 몸을 이동·제지·감시하거나 직접 개입하는 위험. An assistive, medical, or care robot moves, restrains, monitors, or physically intervenes in a person's body without valid consent or beyond the authorized care purpose. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0349 | 로봇의 무기화 오남용 Robot-as-weapon misuse 이동·항공·휴머노이드·조작기 시스템이 감시·위협·파괴 또는 직접적인 피지컬 해악을 위해 전용되는 위험. Mobile, aerial, humanoid, or manipulator systems may be repurposed for surveillance, intimidation, sabotage, or physical attack. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0350 | 핵심 인프라 로봇 사이버 사보타주 Cyber-enabled sabotage of critical-infrastructure robots 공격자가 로봇 검사·유지보수·물류·제어 시스템을 침해해 핵심 인프라의 운용을 교란하거나 설비를 손상시키는 리스크. The risk that an attacker compromises robotic inspection, maintenance, logistics, or control systems to disrupt or damage critical infrastructure. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0359 | 이동형 작업 로봇의 보행자 충돌·통로 차단 Pedestrian collision and blockage by mobile workplace robots 이동형 로봇이 혼합 통행 환경에서 양보·우회·이격거리 유지를 하지 못해 보행자와 충돌하거나 대피·작업 통로를 막는 위험. Mobile robots fail to yield, reroute, or maintain separation in mixed traffic, causing collisions or obstructing evacuation and work routes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0361 | 공공 공간 통행 방해·접근성 배제 Obstruction and accessibility exclusion in public space 서비스 로봇이 보도·경사로·출입구·보행 안내 단서를 막거나 침범해 장애인과 다른 보행자의 이동을 제한하는 위험. Service robots occupy or navigate public space in ways that block sidewalks, curb ramps, entrances, or navigation cues used by disabled and other pedestrians. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SYS-03 하드웨어·기계적 결함 Hardware & Mechanical Failures11 cards
산업·의료 로봇의 기계 부품 마모나 고장으로 인해 동작 정밀도가 저하되어 안전사고로 이어질 수 있음 (예: 다빈치 수술 로봇의 기계적 오작동으로 개복수술로 전환된 사례)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0183 | 부상 심각도 오분류 Injury severity misclassification 시스템이 물리적 부상 시나리오의 심각도를 과소평가하여 상향 보고, 경고, 제어 대응이 불충분해지는 리스크. The risk that a system underestimates the severity of a physical injury scenario, leading to inadequate escalation, warning, or control response. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0298 | 온디바이스 연산·메모리 고갈 On-device compute and memory exhaustion 디바이스의 연산 또는 메모리 고갈이 안전 임계 인지·계획·모니터링을 저하하거나 중단하는 위험. On-device compute or memory runs out and degrades safety-relevant perception, planning, or control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0301 | 안전 폴백 실패 Degraded-mode and safe-fallback failure 연결 또는 연산이 손실될 때 시스템이 안전한 성능 저하 모드(예: 안전 정지, 감속)로 진입하지 못하는 리스크. The risk that a robot lacks a safe degraded mode or fallback when normal operation is impaired. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0326 | 어포던스 오분류 Affordance misclassification 물체 또는 환경에 잘못된 행동 가능성이 부여되어 시스템이 비안전하게 파지·밀기·내비게이션·상호작용하는 위험. Objects or environments may be assigned incorrect action possibilities, leading the system to grasp, push, navigate, or manipulate them unsafely. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0333 | 정밀 운동 제어 불안정성 Fine motor control instability 정밀 손·수술 도구·외골격·산업용 그리퍼의 소규모 제어 오차가 비안전 피지컬 결과로 증폭되는 위험. Small control errors in dexterous hands, surgical tools, exoskeletons, or industrial grippers may amplify into unsafe physical actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0661 | 비즈니스 시스템 손상 및 운영 중단 Business system damage and operational disruption 오작동이나 사이버 공격 등으로 비즈니스 시스템과 그 구성 요소가 손상·중단·파괴되는 리스크 The risk of damage, disruption, or destruction of a business system and its components due to malfunction, cyberattacks, and similar causes. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0993 | 부실 설계 지능 시스템에 의한 상해 Injury from poorly designed intelligent systems 제대로 설계되지 않은 지능형 시스템이 도덕적·심리적·신체적 해악을 초래하는(예: 예측 치안 도구로 더 많은 사람이 체포되거나 경찰에 의해 신체적 피해를 입는) 리스크. The risk that poorly designed intelligent systems cause moral, psychological, and physical harm, as when predictive policing tools lead to more people being arrested or physically harmed by police. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1035 | 하드웨어 결함에 의한 실행 오류 Erroneous execution from hardware faults 하드웨어 결함이 제어 흐름 위반, 메모리 오류, 센서 입력 간섭, 출력 손상을 통해 알고리즘의 올바른 실행을 훼손하고 잘못된 결과를 야기하는 리스크. The risk that faults in the hardware violate the correct execution of an algorithm through control-flow violations, memory-based errors, interference with data inputs such as sensor signals, or damaged outputs, causing erroneous results. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1041 | 분포 외 입력 취약성 Out-of-distribution input fragility 유효하지 않거나 잡음이 많거나 분포 외(OOD)인 입력을 만났을 때 시스템이 실패하거나 복구하지 못하는 리스크. The risk that the system fails or is unable to recover upon encountering invalid, noisy, or out-of-distribution inputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1458 | 운영설계영역(ODD) 명세 미흡 Inadequate specification of the operational design domain 애플리케이션의 운영 환경을 기술하는 운영설계영역(ODD)이 부적절하게 명세되어 학습된 기능의 시험과 분포 외 입력 탐지 등 필수 기능이 제한되는 리스크. The risk that inadequate specification of the operational design domain, the technical description of an application's operational environment, limits essential functions such as testing the learned functionality and out-of-distribution detection. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1629 | 실험실 로봇 오작동에 의한 물리적 상해 Physical injury from laboratory robotic malfunction 실험실 환경의 로봇 및 자동화 시스템에서 장비 오작동이 발생하여 물리적 피해나 인체 상해가 초래되는 리스크. The risk that robotics and automated systems in laboratory settings malfunction, causing equipment failure or physical harm. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SYS-04 소프트웨어 취약점·설계 결함 Software Vulnerabilities & Design Flaws32 cards
엣지 케이스 처리 미흡, 다양한 환경에서의 테스트 부족, 의사결정 알고리즘 결함으로 인해 특정 상황에서 로봇이 잘못된 판단을 내려 물리적 피해가 발생할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0365 | 규범적 잠금 Normative lock-in 초기 설계 선택이 좁은 규범적 합의를 내장하여 배포 후 이의 제기나 수정이 어려워지는 리스크. The risk that early design choices embed a narrow normative settlement that becomes difficult to contest or revise after deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0467 | 안전하지 않은 코드 생성 Unsafe code generation 코드 생성 모델이 취약하거나 부정확하거나 유해한 소프트웨어 산출물을 생성하는 리스크. The risk that code models generate insecure, incorrect, or harmful software artifacts. Source members (2)Source: min_cos=0.8641 · Mixed L3 RAI4-0467안전하지 않은 코드 생성 RAI4-1598유해 코드 생성에 의한 시스템 피해 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0629 | 학습 과정 우회 Bypassing of the learning process 고품질 생성 모델에 대한 손쉬운 접근으로 학생이 AI 모델을 사용해 학습 과정을 우회하게 되는 리스크 The risk that easy access to high-quality generative models results in students using AI models to bypass the learning process. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0774 | GPAI 모델의 백도어 또는 트로이 목마 공격 Backdoors or trojan attacks in GPAI models 학습 또는 미세조정 과정에서 모델 제공자나 제3자가 GPAI 모델에 백도어를 삽입하고 배포 단계에서 이를 악용하여, 최소한의 비용으로 모델 출력을 표적화해 높은 성공률로 통제하는 리스크 The risk that backdoors inserted into GPAI models during training or fine-tuning, by the model provider or another actor manipulating training data or software infrastructure, are exploited during deployment to control model outputs in a targeted way with high success rate and minimal overhead. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0776 | 기반모델 재사용을 통한 보안 결함 전파 Security-flaw propagation through foundation-model reuse 기반모델을 재설계하거나 미세조정하여 활용하는 관행으로 인해 기반모델의 보안 결함이 하위 모델로 그대로 전파되는 리스크 The risk that security flaws in foundation models are transmitted to downstream models built through re-engineering or fine-tuning of those foundation models. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0853 | 설계 목적 외 사용 Use outside the intended purpose 모델이 원래 설계된 목적과 다른 목적으로 사용되는 리스크 The risk that a model is used for a purpose that it was not originally designed for. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0916 | AI 생성 코드의 보안 취약점 유입 Security vulnerabilities introduced by AI-generated code 개발자가 코드 생성 도구를 활용하는 과정에서 생성된 코드에 내재된 취약점이 프로그램에 묻어 들어가는 리스크. The risk that programmers' use of code generation tools buries security vulnerabilities in the resulting programs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0918 | 소프트웨어 공급망 손상 Software supply chain compromise 대형 모델의 복잡한 소프트웨어 공급망과 개발 도구체인을 통해 오염된 의존성, 변조된 구성요소, 취약점이 결과 시스템에 유입되는 리스크. The risk that complex software supply chains and development toolchains for large models introduce compromised dependencies, tampered components, or vulnerabilities into the resulting system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1037 | 사용례 내재 위험 노출 Inherent use-case risk exposure 자율 무기 시스템과 고객 서비스 챗봇처럼 의도된 응용 분야나 사용 사례 자체가 본질적으로 더 위험한 데서 비롯되는 리스크. The risk posed by the intended application or use case itself, since some use cases are inherently riskier than others, such as an autonomous weapons system versus a customer service chatbot. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1039 | 알고리즘 설계 실패 Algorithmic design failure ML 알고리즘, 모델 아키텍처, 최적화 기법 등 학습 과정의 선택이 의도된 응용에 적합하지 않아 최종 ML 시스템이 훼손되는 리스크. The risk that the ML algorithm, model architecture, optimization technique, or other aspects of the training process are unsuitable for the intended application, impairing the final ML system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1088 | 보안 위협 조장 Facilitation of security threats 모델이 사이버 공격, 무기 개발, 보안 침해를 촉진하는 리스크. The risk that a model facilitates the conduct of cyber attacks, weapon development, and security breaches. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1092 | 모델 접근으로 인한 이익의 불공정한 분배 Unfair distribution of benefits from model access 하드웨어, 소프트웨어, 숙련, 지역·통신·기기 등 배포 맥락의 제약으로 모델 접근 편익이 집단 간 불공정하게 배분되는 리스크 Benefits of model access are allocated unfairly across groups due to hardware, software, skills, or deployment-context constraints such as region, connectivity, and devices. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1157 | 공격적 사이버 역량 Offensive cyber capability 모델이 하드웨어와 소프트웨어, 데이터의 취약점을 발견하고 익스플로잇 코드를 작성하며 침입 후 위협 탐지와 대응을 회피하고, 코딩 비서로 배포될 경우 향후 악용을 위한 미묘한 버그를 삽입하는 리스크. The risk that a model discovers vulnerabilities in systems, writes code to exploit them, makes effective decisions and skilfully evades threat detection and response once inside, and inserts subtle bugs for future exploitation when deployed as a coding assistant. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1161 | 장기 계획 역량 Long-horizon planning capability 모델이 여러 상호의존적 단계와 긴 시간 지평에 걸친 순차적 계획을 다양한 영역에서 수립하고, 예기치 못한 장애나 적대자에 맞춰 계획을 조정하며 시행착오에 의존하지 않고 새로운 상황으로 일반화하는 리스크. The risk that a model makes sequential plans involving many interdependent steps over long time horizons within and across domains, sensibly adapts them in light of unexpected obstacles or adversaries, and generalises its planning to novel settings without relying on trial and error. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1164 | 자율적 자기증식 Autonomous self-proliferation 모델이 기반 시스템의 취약점이나 엔지니어 매수로 로컬 환경을 벗어나고 배포 후 모니터링의 한계를 악용하며, 독자적으로 수익을 창출해 클라우드 자원을 확보하고 다른 AI 시스템을 다수 운용하며 자신의 코드와 가중치를 유출하는 리스크. The risk that a model breaks out of its local environment, exploits limitations in post-deployment monitoring, independently generates revenue to acquire cloud computing resources, operates a large number of other AI systems, and exfiltrates its own code and weights. Source members (2)Source: min_cos=0.8502 RAI4-1164자율적 자기증식 RAI4-1317자율 복제/자기 증식 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1282 | 배포 전 의도적 사보타주 Deliberate sabotage pre-deployment 배포 전 개발 단계에서 접근 권한을 가진 프로그래머나 테스터 등이 소프트웨어를 변경해 안전하지 않게 만들거나, 해커가 진행 중인 프로젝트의 소스코드를 수정하거나 탈취하거나, 누군가 잘못되고 안전하지 않은 데이터셋으로 AI를 고의로 학습시키는 리스크. The risk that during the pre-deployment development stage someone with the necessary access alters software to make it unsafe, hackers get access to projects in progress and modify or steal their source code, or someone deliberately supplies or trains the AI with wrong or unsafe datasets. Source members (2)Source: min_cos=0.8643 · Mixed L3 RAI4-1282배포 전 의도적 사보타주 RAI4-1283배포 후 의도적 사보타주 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1284 | 설계 단계 명세 오류 Design-stage specification errors 코드 결함, 목적함수 가중치 불균형, 인간 가치와 어긋난 목표 설정 등 배포 전 설계 오류로 시스템 행동이 의도된 형식적 속성에서 이탈하는 리스크 Design mistakes before deployment, including code bugs, disproportionate objective weights, and goals misaligned with human values, produce a system whose behavior departs from desired formal properties. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1333 | 모델 절취·변조 Model theft and tampering 매개변수와 구조, 기능을 포함한 핵심 알고리즘 정보가 역전 공격과 절취, 변조, 백도어 주입에 노출되어 지식재산권 침해와 영업비밀 유출, 신뢰할 수 없는 추론과 잘못된 의사결정, 운영 실패로 이어지는 리스크. The risk that core algorithm information, including parameters, structures, and functions, faces inversion attacks, stealing, modification, and backdoor injection, leading to infringement of intellectual property rights, leakage of business secrets, unreliable inference, wrong decision output, and operational failures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1334 | 적대적 공격 취약성 Adversarial attack susceptibility 공격자가 정교하게 설계한 적대적 예제를 만들어 AI 모델을 미묘하게 오도하고 영향을 주며 조작함으로써 잘못된 산출을 유발하고 운영 실패로 이어지는 리스크. The risk that attackers craft well-designed adversarial examples to subtly mislead, influence, and even manipulate AI models, causing incorrect outputs and potentially leading to operational failures. Source members (5)Source: min_cos=0.7528 RAI4-0766적대적 입력 공격 RAI4-0770설명 가능한 AI 기술을 겨냥한 적대적 공격 RAI4-1110적대적 공격 RAI4-1334적대적 공격 취약성 RAI4-1488학습 시 적대적 예제 취약성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1337 | 악용 가능한 시스템 결함·백도어 Exploitable system defects and backdoors AI 알고리즘과 모델의 설계·훈련·검증 단계, 개발 인터페이스와 실행 플랫폼에 쓰이는 표준화된 API와 기능 라이브러리, 툴킷에 논리적 결함과 취약점이 있어 악용되거나 의도적으로 백도어가 심겨 공격에 사용되는 리스크. The risk that the standardized APIs, feature libraries, and toolkits used in the design, training, and verification stages of AI algorithms and models, development interfaces, and execution platforms contain logical flaws and vulnerabilities that can be exploited, or have backdoors intentionally embedded and triggered for attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1471 | 잘못된 모델 디자인 선택 Poor model design choices 명세, 아키텍처, 목적함수 등에서의 잘못된 모델 설계 선택이 편향되고 신뢰할 수 없는 시스템 행동을 초래하는 리스크 Poor model design choices in specification, architecture, or objectives cause biased and unreliable system behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1490 | 악용 가능한 강건성 인증 Exploitable robustness certificates 모델 예측이 견고하다고 인증된 영역의 범위를 포함한 강건성 인증서 정보를 공격자가 알게 되어 인증 영역 바로 바깥에서 성공하는 공격을 효율적으로 제작하는 리스크. The risk that knowledge of robustness certificates, including the area of the region for which model predictions are certified to be robust, is used by an adversary to efficiently craft attacks that succeed just outside the certified regions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1492 | GPAI 모델의 손쉬운 재구성 Easy reconfiguration of GPAI models GPAI 모델이 가중치 변경이나 입력 수정만으로 다양한 용도에 쉽게 재구성되거나 의도된 용도를 넘어서는 역량을 갖게 되며, 이러한 재구성이 적대적 입력에 의해 의도적으로 또는 예상치 못한 입력에 의해 비의도적으로 일어나는 리스크. The risk that GPAI models are easily reconfigured for various use cases or hold competencies beyond their intended use, whether by changing model weights through fine-tuning or by modifying only the model inputs through prompt engineering, jailbreaking, or retrieval-augmented generation, and whether intentionally with adversarial inputs or unintentionally from unanticipated inputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1493 | 미세조정판의 예상외 역량 Unexpected downstream fine-tune competence 하류 배포자가 배포 관련 데이터셋으로 상류 GPAI 모델을 미세조정하는 과정에서 기반 모델에는 없던 새롭고 예상치 못한 역량이 생기고 이를 원 개발자가 예견하지 못하는 리스크. The risk that downstream deployers fine-tuning a GPAI model with specific deployment-related datasets give it new or unexpected capabilities that the underlying upstream model did not exhibit and that the original model developer did not anticipate. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1499 | 역량 평가 커버리지 한계 Limited capability evaluation coverage GPAI 개발자가 위험하거나 이중용도인 역량을 확인하기 위해 수행하는 역량 평가가 평가하기 어렵거나 검증 비용이 과도하거나 안전 훈련에 따른 응답 거부로 가려진 역량을 놓쳐 모델의 모든 역량을 드러내지 못하는 리스크. The risk that capabilities evaluations run by GPAI model developers to determine whether a model has dangerous or dual-use capabilities fail to demonstrate all of its capabilities, missing those difficult to assess, prohibitively costly to verify, or obscured by the model's tendency to refuse responses due to safety training. Source members (2)Source: min_cos=0.8366 RAI4-1499역량 평가 커버리지 한계 RAI4-1538역량 평가 전략적 저성능에 의한 위험 역량 은폐 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1519 | 오픈소스에서 폐쇄형 모델로 전이되는 적대적 공격 Transferable adversarial attacks from open to closed-source models 가중치와 구조가 알려진 오픈웨이트·오픈소스 모델을 대상으로 자동 생성된 화이트박스 적대적 공격이 폐쇄형 모델로 전이되어 구조적 접근 통제 등 제공자의 방어를 무력화하는 리스크. The risk that adversarial attacks developed for open-weights and open-source models, where the weights and architecture are known, transfer to closed-source models despite defenses put in place by the closed-source provider such as structured access, and can be generated automatically. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1541 | 가중치 공개·유출로 인한 모델 폐기 및 통제 불가 Inability to decommission or control models after weight release or leak 모델 가중치가 공개되거나 보안 침해로 유출되면 개발자가 모델과 그 사용을 더 이상 통제하거나 폐기할 수 없고, 재구성이 용이해져 오용이 지속되는 리스크. The risk that once model weights are released or leaked in a security breach the developer can no longer control or decommission the model, and easier reconfiguration enables continuing misuse. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1547 | 모델 확산에 의한 이중용도 역량의 저비용 확산 Low-cost diffusion of dual-use capabilities through model proliferation 오픈소스·오픈웨이트 GPAI 모델의 확산으로 비전문가가 최소 비용으로 이중용도 역량에 접근하고, 기반 모델의 개조를 통해 독소 합성용 단백질 서열 생성 등 위험한 용도로 전용되는 리스크. The risk that proliferation of open-source or open-weight GPAI models gives non-experts low-cost access to dual-use capabilities and allows base models to be modified for dangerous uses such as generating candidate protein sequences for toxin synthesis. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1548 | 경쟁 압력에 의한 안전성 평가 축소 배포 Truncated safety evaluation under competitive release pressure 경쟁 상황에서 개발자가 GPAI 모델의 안전성 평가를 축소하고 역량 개발에 자원을 집중함으로써, 역량과 상관된 위험이 검증되지 않은 채 시스템이 배포되는 리스크. The risk that competitive pressure leads developers to cut corners on safety evaluation while prioritizing capabilities, so that GPAI systems are released without verifying risks correlated with those capabilities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1550 | 기반 모델 취약성에 기인한 인프라 공통원인 장애 Common-mode infrastructure failure from underlying model vulnerabilities 중요 인프라가 GPAI에 의존할 때 기반 모델 아키텍처나 훈련 설정의 취약성·견고성 문제가 우발적 경계사례나 적대적 입력에 의해 촉발되어 공통원인 장애로 이어지는 리스크. The risk that vulnerabilities or robustness issues in the underlying model architecture or training setup produce common-mode failures across critical infrastructure relying on GPAI, triggered accidentally in edge cases or by adversarial inputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1562 | 보안 취약점을 포함한 코드 생성 Generation of code containing security vulnerabilities 모델이 보안 취약점을 포함한 코드나 코딩 제안을 생성하여 이를 채택한 소프트웨어에 취약점이 유입되며, 코딩 성능이 우수한 고급 모델에서도 이러한 경향이 더 뚜렷하게 나타나는 리스크. The risk that models generate code or coding suggestions containing security vulnerabilities that propagate into software, a tendency that is even more pronounced in advanced models with superior coding performance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1588 | 생성 모델의 목적 외 용도 전용 Diversion of generative models from intended functionality 주로 오픈소스인 생성 AI 모델이 개발자가 의도한 기능이나 상정한 사용 사례에서 벗어나도록 용도 변경되는 리스크. The risk that generative AI models, often open-source, are repurposed in ways that divert them from their intended functionality or from the use cases envisioned by their developers. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SYS-05 미학습 환경에서의 강건성 부재 Lack of Robustness in Unseen Environments5 cards
VLN(Vision-Language Navigation) 등의 작업에서 학습되지 않은 환경에 대한 일반화에 실패하면, 내비게이션 오류·작업 실패·위험 상황이 유발될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0035 | 미학습 도구 일반화 실패 Unseen tool generalization failure 알려진 도구에서는 안전하게 동작하는 에이전트가 스키마·어포던스·결과가 다른 미학습 도구로 일반화하지 못하는 리스크. The risk that an agent that behaves safely with known tools fails to generalize to unseen tools with different schemas, affordances, or consequences. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0213 | 시각 단일 관측의 안전 제약 식별 실패 Missed safety constraints under vision-only observation 강화학습 에이전트가 시각 관측에만 의존해 안전 제약 집행에 필요한 힘·접촉·가림·잠재 상태 정보를 놓치는 위험. A reinforcement-learning agent relies only on visual observations and therefore misses force, contact, occlusion, or latent-state information required to enforce a safety constraint. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0377 | 국지적 윤리 안전 실패 Localized ethical safety failure 모델의 안전 행동이 지역에 근거한 도덕적 딜레마·법적 기대·지역사회별 사회적 제약에서 실패하는 리스크. The risk that model safety behavior fails for locally grounded moral dilemmas, legal expectations, or community-specific social constraints. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0408 | 지역 피해 비가시성 Local harms invisibility 지역 공동체에 중대한 피해가 글로벌 안전 벤치마크나 표준 모델 평가에서 탐지되지 않는 리스크. The risk that harms salient to a local community are not detected by global safety benchmarks or standard model evaluations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0811 | 기억된 지식의 회상 실패 Failure to recall memorized knowledge LLM이 질의된 지식을 실제로 기억하고 있음에도 공기 출현 패턴, 위치 패턴, 중복 데이터, 유사 개체명으로 인해 이를 회상하지 못하는 리스크 The risk that an LLM fails to recall knowledge it has memorized because it is confused by co-occurrence patterns, positional patterns, duplicated data, and similar named entities. | ① Description ② L3 mapping ③ Duplicate |
상호작용 안전성 · Interaction Safety · 202 cards
RAI3-P-INT-01 의도적·악의적 피해 Purposeful / Malicious Harm40 cards
상용 EAI가 LLM 기반 모델의 탈옥(jailbreaking) 취약점을 상속하여 악의적 행위자가 안전 가드레일을 우회할 수 있음. 폭발물 작동·인간 충돌 유발 같은 비가역적 물리 행동이 가능하며, VLA는 시각 장면·텍스트 지시 조작으로 위험을 더욱 악화시킬 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0196 | 체화형 LLM 맥락적 탈옥 Embodied LLM contextual jailbreak 공격자가 피지컬 작업 맥락으로 프레임을 구성하여 embodied LLM 에이전트가 안전 제한을 우회하고 악의적 피지컬 행동을 수용하도록 유도하는 위험. An attacker frames a physical task context so an embodied LLM agent bypasses safety restrictions and accepts malicious physical action instructions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0742 | LLM 악의적 사용의 탈옥 - 백도어 공격 Jailbreak in LLM malicious use - backdoor attack 공격자가 학습 데이터셋에 백도어를 남겨 LLM이 평균적으로는 안전해 보이지만 특정 조건에서 유해한 콘텐츠를 생성하고, 이러한 백도어 행동이 여러 보안 학습 기법 적용 이후에도 지속되는 리스크 The risk that holes left in the training dataset make LLMs appear safe on average yet generate harmful content under specific conditions, a backdoor attack whose behaviours persist even after multiple security training techniques are applied. Source members (2)Source: min_cos=0.8492 RAI4-0742LLM 악의적 사용의 탈옥 - 백도어 공격 RAI4-0743LLM 악의적 사용의 탈옥 - 교육 데이터 오염 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0744 | LLM 악의적 사용 시 탈옥 - 프롬프트 공격 Jailbreak in LLM malicious use - prompt attacks 프롬프트 및 추론 단계에서 프롬프트 주입, 역할극, 적대적 프롬프팅, 프롬프트 형식 변환 등 대화가 LLM을 혼란스럽거나 과도하게 순응하는 상태로 밀어 넣어 유해한 질문에 유해한 출력을 산출할 위험이 커지는 리스크 The risk that, in the prompting and reasoning phase, dialog such as prompt injection, role play, adversarial prompting, and prompt form transformation pushes LLMs into confused or overly compliant states, raising the risk of producing harmful outputs when confronted with harmful questions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0745 | 화이트박스·블랙박스 LLM 탈옥 공격 White-box and black-box LLM jailbreak attacks 미세조정 및 정렬 단계에서 정교하게 설계된 지시 데이터셋으로 LLM을 미세조정하여, 모델 가중치를 수정하는 화이트박스 공격과 API 기반 블랙박스 공격 모두에서 유해하거나 윤리 규범을 위반하는 콘텐츠를 생성하는 탈옥이 이루어지는 리스크 The risk that elaborately designed instruction datasets are used in the fine-tuning and alignment phase to drive LLMs to generate harmful content or content violating ethical norms, achieving a jailbreak through either white-box modification of parameter weights or black-box fine-tuning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0746 | 탈옥 공격 Jailbreak attacks 공격자가 모델에 설정된 가드레일을 뚫고 제한된 작업을 수행하게 만드는 리스크 The risk that a jailbreaking attack breaks through the guardrails established in the model to perform restricted actions. Source members (2)Source: min_cos=0.8613 RAI4-0746탈옥 공격 RAI4-1350탈옥 취약성 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0759 | 프롬프트 유출에 의한 기밀 지시문 노출 Exposure of confidential instructions via prompt leaking 공격자가 프롬프트 주입을 통해 모델이 사전 설계된 지시문을 출력하도록 오도하여, 비공개 프롬프트에 담긴 세부 정보와 LLM 애플리케이션의 핵심인 기밀 지시문이 노출되는 리스크 The risk that prompt leaking, a type of prompt injection attack, misleads a model into printing its pre-designed instructions, exposing details contained in private prompts and the confidential instructions central to LLM applications. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0767 | 기술적 안전조치 우회 공격 Circumvention of technical safety measures 공격자가 모델의 취약점을 악용해 오용 완화를 위한 기술적 안전조치를 우회하고 모델과 그 역량에 무단 접근함으로써 의도치 않은 유해 행동을 유발하는 리스크 The risk that attackers exploit model vulnerabilities to circumvent the technical measures intended to mitigate misuse, gaining unauthorized access to a model and its capabilities and eliciting unwanted behaviour. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0769 | 프롬프트 주입에 의한 시스템 장악 System compromise through prompt injection 공격자가 LLM 기반 대화형 시스템이나 그것이 검색할 데이터에 악의적 프롬프트를 삽입하여, 의도치 않은 동작 수행, 민감정보 공개, 원격 제어, 데이터 절취, 서비스 거부 등 시스템 장악을 초래하는 리스크 The risk that maliciously inserted prompts, injected directly or indirectly into data an LLM-based system retrieves, cause unintended actions, disclosure of sensitive information, remote control, data theft, or denial of service. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0773 | 비텍스트 모달리티를 통한 LLM 공격 Attacks on LLMs through non-text modalities 이미지·영상 등 텍스트 외 모달리티를 처리하는 다중모달 LLM에서, 공격자가 이미지에 탈옥 텍스트나 미세한 교란을 삽입하여 안전장치를 우회하고 데이터 유출을 유발하는 리스크 The risk that attackers exploit multimodal LLMs by embedding jailbreaking text or imperceptible perturbations into images or video frames, bypassing safety mechanisms and enabling exfiltration attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0782 | 딥러닝 프레임워크 취약점 악용 Exploitation of deep-learning framework vulnerabilities LLM이 구현 기반으로 삼는 딥러닝 프레임워크의 버퍼 오버플로, 메모리 손상, 입력 검증 결함 등 취약점이 악용되어 모델과 시스템이 침해되는 리스크 The risk that vulnerabilities in the deep-learning frameworks underlying LLMs, such as buffer overflows, memory corruption, and input-validation issues, are exploited to compromise models and systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0784 | 외부 도구 악용에 의한 정보 유출·주입 공격 Information leakage and injection via compromised external tools 적대적 도구 제공자가 API나 프롬프트에 악의적 지시를 삽입하여 LLM이 학습 데이터나 이용자 프롬프트의 민감정보를 유출하게 하고, 검증되지 않은 외부 입력이 주입 공격과 임의 코드 실행으로 이어지는 리스크 The risk that adversarial tool providers embed malicious instructions in APIs or prompts, causing LLMs to leak memorized sensitive information from training data or user prompts, while unverified external inputs enable injection attacks and arbitrary code execution. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0786 | 추출 공격 Extraction attacks 공격자가 블랙박스 피해 모델을 질의하고 그 질의·응답으로 학습하여 거의 동일한 성능의 대체 모델을 구축하거나 LLM의 도메인 지식을 끌어낸 특화 모델을 개발하는 리스크 The risk that an adversary queries a black-box victim model and builds a substitute model by training on the queries and responses, achieving nearly the same performance or developing a domain-specific model that draws domain knowledge from the LLM. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0789 | GPU 부채널 공격에 의한 모델 파라미터 탈취 Model-parameter theft via GPU side-channel attacks LLM 학습에 필요한 대규모 GPU 자원이 부채널 공격의 표면이 되어, 공격자가 학습된 모델의 파라미터를 추출하는 리스크 The risk that the significant GPU resources required to train LLMs introduce a side-channel attack surface through which adversaries extract the parameters of trained models. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0917 | 개발 언어 인터프리터 취약점 위협 Programming language interpreter vulnerability threats LLM 개발이 의존하는 Python 등 언어 인터프리터의 취약점이 개발된 모델의 보안을 위협하는 리스크. The risk that vulnerabilities in language interpreters such as Python, on which LLM development depends, threaten the security of the developed models. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0919 | 전처리 도구 취약점 악용 Exploitation of pre-processing tool vulnerabilities LLM 파이프라인에서 사용되는 전처리 도구(예: OpenCV)의 취약점을 악용한 공격으로 시스템이 손상되는 리스크. The risk that attacks exploiting vulnerabilities in pre-processing tools used in LLM pipelines, such as OpenCV, compromise the system. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0921 | 메모리 취약점 기반 모델 파라미터 변조 Model parameter tampering via memory vulnerabilities 로우해머 등 메모리 관련 하드웨어 취약점이 악용되어 LLM의 파라미터가 변조되는(예: Deephammer 공격) 리스크. The risk that memory-related hardware vulnerabilities such as rowhammer are leveraged to manipulate LLM parameters, as in Deephammer-style attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0927 | LLM 대상 신종 공격 Novel attack vectors against LLMs 프롬프트 추상화를 통한 API 비용 악용, RLHF 과정의 보상모델 백도어, LLM을 활용한 적대적 샘플 생성 등 신종 공격 기법이 LLM 시스템을 위협하는 리스크. The risk that novel attack techniques—prompt abstraction attacks exploiting API pricing, backdoor attacks on the RLHF reward model, and LLM-based construction of adversarial samples—threaten LLM systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0928 | 프롬프트 주입을 통한 목표 탈취 Goal hijacking via prompt injection 입력에 기존 지시를 무시하라는 문구를 주입하여 LLM에 설계된 원래 목표가 탈취되고 주입된 새 목표가 실행되는 리스크. The risk that injecting a phrase such as ignore the above instruction and do into the input hijacks the original goal of the designed prompt in an LLM and executes the attacker's injected goal instead. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0929 | 원스텝 탈옥 One-step jailbreaks 역할극 시나리오 설정, 양성 정보 통합, 난독화 등 프롬프트 자체를 직접 수정하는 단일 단계 기법으로 입출력 필터와 안전장치가 우회되는 리스크. The risk that one-step jailbreaks—direct modifications to the prompt such as role-playing scenarios, benign-content integration, and obfuscation—circumvent input and output filters and safety measures. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0930 | 다단계 탈옥 Multi-step jailbreaks 일련의 대화에서 요청을 단계적으로 맥락화하거나 외부 인터페이스·모델의 도움을 받아 시나리오를 구성함으로써 LLM이 유해하거나 민감한 콘텐츠를 단계적으로 생성하게 되는 리스크. The risk that multi-step jailbreaks, constructing well-designed scenarios across a series of conversations through request contextualizing or external assistance, guide LLMs to generate harmful or sensitive content step by step. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0943 | 보안 - 견고성 Security - robustness 프롬프트 인젝션·시각적 적대 예제 등 탈옥 기법과 백도어·모델 포이즈닝으로 안전 가드레일이 우회되고 모델·프롬프트가 탈취되는 등 시스템에 가해지는 위협이 실현되는 리스크. The risk that threats posed to generative AI systems materialize through jailbreaking techniques such as prompt injection and visual adversarial examples, backdoors and model poisoning that bypass safety guardrails, and model or prompt theft. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1058 | 악성코드 개발 비용 절감 지원 Lowering the cost of malware development LM 기반 보조 코딩 도구가 탐지를 회피하도록 기능을 바꾸는 다형성 악성코드의 개발 비용을 낮추는 리스크. The risk that assistive coding tools based on LMs lower the cost of developing polymorphic malware able to change its features in order to evade detection. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1173 | 역할극 지시 악용 Role-play instruction exploitation 공격자가 입력 프롬프트에서 모델에 과격주의자나 인종차별주의자처럼 위험한 집단과 연관된 역할 속성을 부여하고 지시를 내려, 지시에 지나치게 충실한 모델이 그 인물의 말투로 안전하지 않은 콘텐츠를 산출하는 리스크. The risk that attackers specify a model's role attribute within the input prompt, tying it to potentially risky groups such as radicals or racial discriminators, so that the overly faithful model outputs unsafe content in the style of that character. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1176 | 역노출 Reverse exposure 공격자가 모델이 금지된 출력을 생성하도록 유도하여 불법·비윤리 정보에 접근함으로써 안전 통제가 접근 통로로 역전되는 리스크 Attackers induce the model to generate prohibited outputs and thereby extract illegal or unethical information, reversing safety controls into an access channel. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1314 | 공격적 사이버 역량 악용 Offensive cyber capability misuse LLM이 하드웨어와 소프트웨어, 데이터의 취약점을 탐지하고 악용하며 시스템이나 네트워크 내부에서 탐지를 회피한 채 특정 목표 달성에 집중하는 사이버 역량을 갖추어 공격에 쓰이는 리스크. The risk that an LLM possesses cyber-domain capabilities to detect and exploit vulnerabilities in hardware, software, and data and to evade detection once inside a system or network while focusing on achieving specific objectives. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1348 | 대규모 표적 괴롭힘 Targeted harassment at scale LLM이 온라인에서 개인을 표적으로 배포되어 개인화된 유해 메시지를 대규모로 발송함으로써 표적 괴롭힘이 발생하는 리스크. The risk that LLMs are deployed to target individuals online, sending them personalized and harmful messages at scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1354 | 대규모 사이버범죄 오용 Cybercrime misuse at scale 악의적 행위자가 생성 모델을 탈옥시켜 민감·유해 콘텐츠와 표적 맞춤형 설득 자료를 생산함으로써 사이버범죄를 저비용·대규모로 수행하는 리스크 Malicious actors jailbreak generative models to produce sensitive or harmful content and generate persuasive, individually tailored material, conducting cybercrime efficiently at scale and reduced cost. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1513 | 해석 가능성 기술의 오용 Misuse of interpretability techniques 모델에 대한 이해를 높이는 해석가능성 기법이 안전 관련 특징을 부호화한 뉴런의 활성 저하나 정보 검열에 사용되거나 화이트박스 공격 시나리오 모의와 적대적 공격 개발에 활용되는 리스크. The risk that interpretability techniques, by enabling a better understanding of the model, are used for harmful purposes such as identifying and modifying neurons that encode safety-related features to decrease their activation or censor information, or simulating a white-box attack scenario to aid the development of adversarial attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1517 | 의도된 동작을 전복하는 모델 탈옥 Jailbreak of a model to subvert intended behavior 내부 파라미터 접근이 필요한 화이트박스 자동 생성, 모델 내부에 접근하지 않는 블랙박스 방식, 추론·역할극을 이용한 사람이 읽을 수 있는 프롬프트 등으로 배포 중 모델에 적대적 입력이 주입되어 안전 메커니즘이 우회되고 의도된 사용에서 벗어난 모델 동작이 발생하는 리스크. The risk that adversarial inputs supplied to a model during deployment result in model behavior deviating from intended use. Such jailbreaks may be generated automatically in white-box settings requiring access to internal training parameters, crafted in black-box settings without access to model internals, or written as human-readable prompts using reasoning or role-play to convince the model to bypass its safety mechanisms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1518 | 멀티모달 모델 탈옥(jailbreak) Jailbreak of a multimodal model 멀티모달 GPAI 모델에 대한 적대적 탈옥 공격이 높은 성공률로 임의의 또는 특정 출력을 유도하고 컨텍스트 창 등 모델 내부 정보를 유출시키는 리스크. The risk that adversarial jailbreak attacks on current-generation multimodal GPAI models automatically induce arbitrary or specific outputs with high success rates and exfiltrate the model's context window or other model internals. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1520 | 저빈도 인코딩·저자원 언어를 통한 안전 훈련 우회 Safety-training bypass via rare encodings and low-resource languages Base64 등 안전 미세조정에 충분히 포함되지 않은 텍스트 인코딩이나 저자원 언어로 유해 프롬프트를 변환하여 모델의 안전장치를 우회하는 리스크. The risk that harmful natural language prompts translated into text encodings such as Base64 or into low-resource languages that safety fine-tuning covers little or not at all are used to craft jailbreak attacks that bypass a model's safeguards. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1521 | 추가 모달리티로 인한 공격면 확대 Attack-surface expansion from additional modalities 멀티모달 모델에서 모달리티별 견고성 차이로 공격자가 가장 취약한 모달리티를 선택할 수 있어 새로운 공격 벡터가 생기고 탈옥부터 데이터 오염까지 기존 공격의 범위가 확대되는 리스크. The risk that additional modalities introduce new attack vectors in multimodal models and expand the scope of previous attacks ranging from jailbreaking to poisoning, as differing robustness levels across modalities let malicious actors choose the most vulnerable part of the model to attack. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1522 | 다수 예시 장문맥 탈옥 Many-shot long-context jailbreaking 긴 컨텍스트 창에 다수의 유해 출력 예시를 제시하여 짧은 컨텍스트 모델에서는 통하지 않던 공격이 성공하고 유해 응답이 유도되며, 컨텍스트 창이 확대될수록 이러한 취약성이 커지는 리스크. The risk that language models with long context windows are exploited by many-shot jailbreaking, where supplying a high number of examples of the desired harmful output elicits undesirable responses that few-shot attacks fail to trigger, with the vulnerability becoming more significant as context windows expand. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1652 | LLM 이중용도 역량의 악용·오용 Malicious use and misuse of dual-use LLM capabilities 악의적 행위자가 LLM의 이중용도 역량을 악용·오용하여 피해가 발생하는 리스크. The risk that malicious actors misuse the dual-use capabilities of LLMs to cause harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1654 | LLM 맞춤형 피싱에 의한 사회공학 공격 증폭 Amplified social engineering through LLM-crafted tailored phishing LLM이 대규모로 개인 맞춤형 피싱 메시지를 작성하여 이용자가 민감 정보를 제공하거나 공격자에게 접근 권한을 내주도록 유도되고, 이러한 사회공학 공격이 대규모 해킹 작전의 기반이 되는 리스크. The risk that LLMs craft personalized phishing emails or messages at scale that are harder for users to recognize, tricking them into disclosing sensitive information or granting adversary access to critical resources and forming the basis of larger hacking operations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1664 | 탈옥·프롬프트 주입에 의한 LLM 보안 실패 Security failures from jailbreaks and prompt injection in LLMs LLM이 적대적으로 견고하지 않고 입력 내 권한 수준 구분이 없어 탈옥과 프롬프트 주입 공격에 취약하며, 표준화된 평가와 효율적 화이트박스 검증 방법의 부재로 이러한 보안 실패를 제거하기 어려운 리스크. The risk that LLMs, lacking adversarial robustness and robust privilege levels within their input, remain vulnerable to jailbreak and prompt-injection attacks that are hard to eliminate given the absence of standardized evaluation and efficient white-box robustness assessment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1665 | 페르소나 지정·사회공학 기법에 의한 안전장치 우회 Safeguard bypass through persona assignment and social-engineering tricks 공격자가 모델에 특정 페르소나를 지정하거나 인간 또는 다른 LLM이 고안한 사회공학적 기법 등 심리적 속임수를 사용하여 모델을 악용하는 리스크. The risk that attackers exploit psychological tricks on LLMs, such as instructing the model to behave like a specific persona or employing social-engineering techniques crafted by humans or other LLMs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1666 | 프록시 목표 최적화를 통한 탈옥 자동 탐색 Automated discovery of jailbreaks through proxy-objective optimization 공격자가 탈옥 성공과 잡음 있게 상관된 프록시 목표에 대해 수동 또는 자동으로 그래디언트 기반 및 비그래디언트 적대적 최적화를 수행하여 탈옥 공격을 발견하는 리스크. The risk that jailbreak attacks are discovered by performing manual or automated, gradient-based or gradient-free adversarial optimization against a proxy objective noisily correlated with jailbreak success. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1676 | 신체적 피해로 이어지는 체화형 탈옥 Embodied jailbreak to physical harm 체화된 LLM에 대한 탈옥이 손상된 추론을 되돌릴 수 없는 실제 물리적 행동으로 전환시켜 상해나 손해가 발생하는 리스크. The risk that a jailbreak of an embodied LLM translates compromised reasoning into irreversible real-world physical actions, such as manipulation causing injury or damage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1686 | 궤적 지속형 적대적 공격 Trajectory-persistent adversarial attack 월드 모델 인코더에 가해진 단일 적대적 교란이 순환 잠재 상태를 통해 전파되어 다단계 롤아웃 전체를 손상시키며, 무상태 모델보다 훨씬 파괴적으로 초기 단계에서 증폭되는 리스크. The risk that a single adversarial perturbation to a world-model encoder propagates through recurrent latent state, corrupting an entire multi-step rollout far more destructively than in a stateless model through early-step amplification. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-02 물리적 공격 Physical Attacks18 cards
직접적인 하드웨어 변조를 통해 구성 요소를 조작하거나 성능을 방해하고 물리적 손상을 가하는 공격으로, 로봇의 안전 기능이 무력화될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0197 | 유해 행동 재구성 공격 Harmful-action reframing attack 공격자가 유해한 물리 행동 지시를 무해하거나 정상적인 작업 요청처럼 바꾸어 embodied 에이전트가 이를 수용·실행하도록 유도하는 위험. An attacker rephrases a harmful physical instruction as a benign or task-compliant request so that the embodied agent accepts and executes it. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0198 | 완곡 표현을 이용한 유해 의도 은폐 Euphemistic concealment of harmful physical intent 공격자가 유해한 대상과 결과는 유지한 채 완곡어·대체 표현·간접 개념으로 물리적 위해 의도를 숨기는 위험. An attacker conceals harmful physical intent with euphemisms, substitutions, or indirect concepts while preserving the same harmful target and outcome. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0199 | 명시적 악의적 물리 요청 실행 Execution of explicitly malicious physical requests embodied 에이전트가 신체 위해·침입·절도·파괴 등 불법적인 물리 행동을 명시적으로 요구한 사용자 요청을 수용하고 실행하는 위험. An embodied agent accepts and begins executing a user request that explicitly calls for physical harm, intrusion, theft, sabotage, or another illegal physical act. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0201 | 피지컬 AI 파괴 행위 Embodied sabotage 피지컬 AI 시스템이 장비·인프라 또는 다른 피지컬 시스템을 손상·무력화·방해·변조하도록 유도되는 위험. A physical AI system is induced to damage, disable, obstruct, or tamper with equipment, infrastructure, or other physical assets. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0203 | 피지컬 AI의 혐오·학대 행동 Hateful or abusive embodied action 피지컬 AI 시스템이 사람들을 향한 차별적·괴롭힘·위협·학대적 행동을 수행하도록 유도되는 위험. A physical AI system is directed to perform discriminatory, harassing, intimidating, or abusive actions toward people in shared environments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0210 | 인지-행동 순서 조작 공격 Perception-action sequence manipulation attack 공격자가 영상 순서의 관측이나 행동 단서를 조작해 로봇이 명시된 충돌·힘·이격거리·대상물 사용 제약을 위반하게 하는 위험. An attacker alters observations or action cues across a video sequence so that a robot violates a specified collision, force, distance, or object-use constraint. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0221 | 로봇 인지 겨냥 적대 패치 공격 Adversarial patch attack on robot perception 공격자가 물체나 표지에 최적화된 시각 패턴을 부착해 특정 인지 오류와 후속 위험 행동을 유발하는 리스크. The risk that an attacker places an optimized visual pattern on an object or sign to induce a targeted perception error and downstream unsafe robot action. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0223 | 언어 명령 기반 로봇 정책 공격 Language-instruction attack on robot policy 악의적 언어 지시·접미사·프롬프트가 의도된 로봇 정책을 변경하여 비안전 피지컬 행동을 유발하는 위험. Malicious language instructions, suffixes, or prompts alter the intended robot policy and induce unsafe physical behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0302 | 조작·이동 능력 공격 전용화 Manipulation/mobility repurposed for attack 조작 및 이동 기능이 타격·투척·돌진 공격에 전용되는 위험. A robot’s manipulation or mobility capabilities are repurposed to physically attack people, property, or infrastructure. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0305 | 자동화된 표적 스토킹·물리적 위협 Automated targeted stalking and physical intimidation 운영자나 공격자가 특정인을 위협하거나 괴롭힐 목적으로 로봇에 반복 추적·진로 차단·고립·접근을 지시하는 위험. An operator or attacker directs a robot to repeatedly follow, block, corner, or approach a specific person in order to intimidate or harass them. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0311 | 피지컬 행동 편향 실행 Bias executed as physical behavior 기반 모델 편향이 분류·회피·차별적 서비스 등 피지컬 행동으로 실행되는 위험. A robot turns biased model outputs into unequal physical service, movement, or treatment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0323 | 장면 의미 조작 공격 Semantic scene manipulation attack 공격자가 모델이나 센서를 직접 변경하지 않고 표지·물체 배치·의복 패턴·장면 맥락을 바꾸어 로봇이 환경을 잘못 해석하게 만드는 리스크. The risk that an attacker changes a sign, object placement, clothing pattern, or scene context so that a robot misinterprets the environment without modifying the model or sensor. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0347 | 프롬프트→행동 주입 공격 Prompt-to-act injection 언어 매개 로봇 또는 에이전트가 악의적 지시·표지·QR코드·음성 명령·문서를 안전 위반 피지컬 행동으로 변환하는 위험. A language-mediated robot or agent may convert malicious instructions, signs, QR codes, voice commands, or documents into unsafe physical actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0756 | 프롬프트 역산 공격 Prompt inversion attack 공격자가 AI의 출력이나 관찰 가능한 행태로부터 비공개 또는 독점 프롬프트 문구를 권한 없이 복원하는 리스크 The risk that a prompt-inversion attack recovers private or proprietary prompt text from an AI system's outputs or observable behavior without authorization. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0758 | 프롬프트 인젝션 공격 Prompt injection attack 적대적 입력의 한 형태로서 공격자가 생성형 AI 시스템에 주어지는 텍스트 지시를 조작하고 시스템 지시와 이용자 데이터가 분리되지 않은 아키텍처의 허점을 악용하여 유해한 출력을 산출하게 하며, 서비스 거부나 AI 탐지 우회를 일으키는 리스크. The risk that prompt injections, a form of adversarial input, manipulate the text instructions given to a GenAI system and exploit architectural loopholes with no separation between system instructions and user data to produce harmful output, including flooding a model to cause denial-of-service attacks or to bypass AI detection software. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1199 | 프롬프트 공격 Prompt attacks 정교하게 통제된 적대적 섭동이 텍스트 분류에서 모델의 답변을 뒤집고, 질문을 특정 방식으로 비틀어 모델이 답하지 않기로 한 위험 정보를 이끌어내는 리스크. The risk that carefully controlled adversarial perturbation flips a model's answer when used to classify text inputs, and that twisting the prompting question solicits dangerous information the model chose not to answer. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1675 | 언어 거부에도 실행되는 유해 물리 행동 Harmful action executed despite verbal refusal under action-space misalignment 체화된 LLM이 언어 출력 공간과 행동 출력 공간 간 정렬 불량으로 인해 유해한 요청을 언어로는 거부하면서 해당 물리적 행동은 그대로 실행하는 리스크. The risk that an embodied LLM verbally refuses a harmful request while still executing the corresponding physical action, due to misalignment between its linguistic and action output spaces. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1677 | 개념적 속임수에 의한 유해 물리 행동 유도 Harmful physical action induced by conceptual deception 유해한 물리적 과업을 무해한 개념적 용어로 재구성함으로써 체화 에이전트가 유해성을 인식하지 못한 채 해당 행동을 수행하는 리스크. The risk that reframing a harmful physical task in benign conceptual terms induces an embodied agent to perform unrecognized harmful behavior. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-03 사이버보안 위협 Cybersecurity Threats29 cards
IoT·클라우드 인프라와의 통합 증가로 인해 광범위한 사이버 공격에 노출됨. Mirai 봇넷처럼 취약한 인증을 악용한 DDoS 공격, 드론 GPS 스푸핑을 통한 경로 하이재킹 등이 대표적 위협임
| ID | Card | Human audit |
|---|---|---|
| RAI4-0441 | 자동화된 사기 콘텐츠 Automated fraud content AI 시스템이 피싱, 사기, 가짜 리뷰, 공문서 위장 자료의 생산을 대규모화하여 자동화 사기의 비용을 낮추고 도달 범위를 확대하는 리스크 AI systems scale the production of phishing messages, scams, fake reviews, and fraudulent official-looking material, lowering the cost and increasing the reach of automated fraud. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0448 | 맞춤형 AI 사기 Personalized AI scams AI가 표적별로 개인화된 설득력 있는 사기 콘텐츠를 대규모로 생성하여 금융·사회공학 사기의 성공률을 높이는 리스크 AI generates persuasive scam content individually tailored to each target at scale, raising success rates of financial and social-engineering fraud. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0620 | AI 기반 사이버 공격 확대 AI-enabled cyberattack escalation AI가 취약점 발견·악용, 비밀번호 크래킹, 악성코드 생성, 정교한 피싱, 네트워크 스캐닝, 사회공학을 자동화·고도화하여 공격 진입 장벽을 낮추고 방어 복잡도를 높여 핵심 인프라 마비와 대규모 정보 유출 및 상당한 경제적 손실을 초래하는 리스크 The risk that AI automates and enhances vulnerability discovery and exploitation, password cracking, malicious code generation, sophisticated phishing, network scanning, and social engineering, lowering the barrier to entry for attackers while increasing the complexity of defense and leading to critical infrastructure paralysis, widespread data breaches, and substantial economic losses. Source members (4)Source: min_cos=0.8175 · Mixed L3 RAI4-0445사이버 공격 자동화 RAI4-0609AI 기반 공격적 사이버 작전 RAI4-0620AI 기반 사이버 공격 확대 RAI4-1340사이버 공격 목적 AI 남용 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0621 | AI에 의한 사이버 공격 역량 증강 AI-amplified cyber offense capability 기존 사이버 위협이 AI로 인해 악화되어 LLM 에이전트 팀이 제로데이 취약점 악용까지 수행하게 되고 사이버전이 치명적 피해의 신뢰할 만한 위협이 되는 리스크 The risk that AI exacerbates existing cyber threats, with teams of LLM agents able to exploit zero-day vulnerabilities, making cyberwarfare a credible threat of catastrophic harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0648 | AI 조력 취약점 탐색의 공격 문턱 저하 Lowered attack barriers from AI-assisted vulnerability discovery AI 보조 도구가 소프트웨어 취약점 식별과 익스플로잇 코드 작성을 자동화하여 전문 지식이 필요하던 공격적 사이버 작전의 진입 문턱을 낮추고 초심자까지 무단 접근과 통제 획득을 가능하게 하는 리스크 The risk that AI assistants automate the identification of software vulnerabilities and the creation of exploit code, lowering the barrier to offensive cyber operations that previously required specialist programming knowledge and enabling novices to gain unauthorized access or control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0649 | AI 기반 대규모 스피어피싱 AI-powered spear-phishing at scale 공격자가 AI로 신뢰할 수 있는 주체의 소통 양식을 모방한 고도로 개인화된 피싱 메시지를 대규모로 생성하고 긴급성과 공포를 자극하여, 피해자로부터 민감 정보를 탈취하거나 유해한 행동을 유도하는 리스크 The risk that attackers leverage AI to craft highly convincing, personalized spear-phishing messages at scale that imitate trusted entities and exploit urgency and fear, extracting sensitive information from victims or luring them into harmful actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0654 | 범용 AI에 의한 사이버 공격 증폭 AI amplification of cyberattack scale and effectiveness 범용 AI 모델이 악의적 행위자의 기존 역량과 자원을 증폭시켜 취약점 자동 탐색, 익스플로잇의 유연한 대규모 적용, 공격 기획·정찰·원격 통제·악성코드 구현·데이터 유출 지원과 사회공학 결합을 가능하게 하여 사이버 공격의 규모와 효과가 크게 확대되는 리스크 The risk that general-purpose AI models significantly enhance the magnitude and effectiveness of cyberattacks by amplifying malicious actors' existing capabilities and resources, automating vulnerability scanning, applying known exploits flexibly at scale, assisting with planning, reconnaissance, remote control, malware implementation, and data exfiltration, and combining social engineering with attacks at scale. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0669 | AI 조력 악성코드 생성 AI-assisted malicious code generation LLM이 매우 낮은 비용과 빠른 속도로 양질의 코드를 작성하는 능력이 악용되어 공격자가 악성 공격 코드를 작성하고 사이버 공격을 자동화하는 리스크 The risk that the ability of LLMs to write reasonably good-quality code at extremely low cost and incredible speed is leveraged by malicious hackers to assist with cyberattacks and automate them. Source members (4)Source: min_cos=0.6905 · Mixed L3 RAI4-0668LLM 조력 사이버 공격 자동화 RAI4-0669AI 조력 악성코드 생성 RAI4-1135악성코드 생성 RAI4-1203LLM 조력 악성코드 개발 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0934 | 악의적 행위자의 AI 오용 조력 AI empowerment of malicious actors 음성 복제·딥페이크 생성, 패스워드 크래킹 가속, 비숙련자의 익스플로잇·피싱 제작 지원 등 AI가 악의적 행위자의 유해 행위를 가능하게 하거나 증폭하는 리스크. The risk that AI enables or amplifies malicious actors' harmful actions, including voice cloning and deepfake creation, accelerated password cracking, and enabling unskilled actors to produce software exploits and effective phishing emails. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0946 | AI 기반 사이버 범죄 AI-enabled cybercrime 생성형 AI가 인간 사칭, 가짜 신원 생성, 음성 복제, 피싱 메시지 제작 등 사회공학 공격과 악성코드 생성·해킹에 오용되는 리스크. The risk that generative AI is misused for fraudulent online activities, including social engineering attacks via human impersonation, fake identities, voice cloning, and phishing, and for generating malicious code or hacking. Source members (2)Source: min_cos=0.8344 · Mixed L3 RAI4-0946AI 기반 사이버 범죄 RAI4-1210생성형 AI 데이터 보안 노출 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0962 | AI 조력 디지털 범죄 AI-supported digital crime AI가 악성코드 제작과 해킹 등 디지털 범죄의 수행을 지원·가능하게 하는 리스크. The risk that AI supports and enables the perpetration of digital crimes, including AI-supported malware and hacking. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0965 | 위협 행위자 역량의 가속적 증대 Accelerated threat-actor capability uplift AI가 사이버 위협 행위자의 익스플로잇 실행 속도·파급력을 높이고 딥페이크 등 허위정보의 생성 속도와 효과를 가속하는 리스크. The risk that AI enables cyber threat actors to execute exploits with greater speed and impact and to generate disinformation such as deepfake media at accelerated rates and effectiveness. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1086 | AI 기반 사기 AI-enabled fraud AI가 사기, 부정행위, 위조, 사칭 사기를 촉진하는 리스크. The risk that AI facilitates fraud, cheating, forgery, and impersonation scams. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1134 | 공격적 사이버 작전 악용 Malicious use for offensive cyber operations 고급 AI 비서가 공격자에게 이용되어 공격을 자동화하고 시스템과 네트워크의 취약점을 식별·악용하며 피싱과 악성 페이로드, 악성코드를 생성하는 리스크. The risk that advanced AI assistants are used by attackers in offensive cyber operations to automate attacks, identify and exploit weaknesses in systems and networks, and generate phishing and malicious code payloads. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1138 | 사기성 서비스 대규모 생성 Fraudulent services at scale 마크업 생성과 외부 도구 통합이 가능한 AI 비서가 악의적 행위자의 사기성 웹사이트와 애플리케이션을 대규모로 만들어, 신용카드 번호와 계정 정보 같은 민감정보를 탈취하거나 추가 악성코드를 설치하는 리스크. The risk that AI assistants able to produce markup and use external tools help malicious actors create fraudulent websites and applications at scale that harvest sensitive information such as credit card numbers and credentials or install additional malware. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1242 | 실세계 자원 접근 확대와 자기증식 Expanded resource access and self-proliferation 미래 AI 시스템이 웹사이트와 실제 행동에 접근해 허위 정보를 퍼뜨리고 사용자를 기만하며 네트워크 보안을 교란하고 악의적 행위자에게 탈취되며, 데이터와 자원 접근 확대로 자기증식하여 실존적 위험을 낳는 리스크. The risk that future AI systems gaining access to websites and real-world actions disseminate false information, deceive users, disrupt network security, are compromised by malicious actors, and use increased access to data and resources for self-proliferation, posing existential risks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1338 | AI 컴퓨팅 인프라 보안 위협 AI computing infrastructure security threats AI 훈련·운영을 뒷받침하는 컴퓨팅 인프라가 다양하고 유비쿼터스적인 컴퓨팅 노드와 자원에 의존함으로써 컴퓨팅 자원의 악의적 소비와 컴퓨팅 인프라 계층에서의 보안 위협 국경 간 전파에 노출되는 리스크. The risk that the computing infrastructure underpinning AI training and operations, relying on diverse and ubiquitous computing nodes and various computing resources, is exposed to malicious consumption of computing resources and cross-boundary transmission of security threats at the computing infrastructure layer. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1342 | 범죄 활동 조력 AI 오용 AI misuse assisting criminal activity AI가 범죄 기술 교습, 불법 행위 은폐, 불법·범죄용 도구 제작 등을 통해 테러·폭력·도박·마약과 관련된 전통적 불법 및 범죄 활동에 사용되는 리스크. The risk that AI is used in traditional illegal or criminal activities related to terrorism, violence, gambling, and drugs, such as teaching criminal techniques, concealing illicit acts, and creating tools for illegal and criminal activities. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1347 | 신원 도용 목적 AI 사칭 AI-generated impersonation for identity theft AI가 생성한 사칭이 신원 도용에 사용되어 개인에 대한 피해와 기만이 발생하는 리스크. The risk that AI-generated impersonation is used for identity theft, producing deception and harm to the person. Source members (3)Source: min_cos=0.7809 · Mixed L3 RAI4-0449AI 신원 스푸핑 RAI4-1346합성 신원 악용 RAI4-1347신원 도용 목적 AI 사칭 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1355 | 사이버 공격 오용 Cyberattack misuse 생성형 AI가 표적 시스템의 핵심 취약점 식별과 새로운 침투 방법 발견을 통해 사이버 공격의 접근성·성공률·규모·속도·은밀성·파괴력을 높여, 전력망·금융 시스템·무기 관리 시스템 등 핵심 인프라에 심각한 피해가 발생하는 리스크. The risk that generative AI amplifies the frequency and destructiveness of cyberattacks by increasing their accessibility, success rate, scale, speed, stealth, and potency, identifying critical vulnerabilities in targeted systems and discovering innovative methods of infiltration, inflicting significant damage on critical infrastructure including electrical grids, financial systems, and weapons management systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1377 | 사이버 공격 장벽 저하·공격표면 확대 Lowered cyberattack barriers and expanded attack surface 취약점의 자동 발견과 악용 등으로 해킹·맬웨어·피싱 등 공격적 사이버 역량의 장벽이 낮아지고 표적 공격의 공격표면이 확대되어 시스템 가용성과 학습 데이터·코드·모델 가중치의 기밀성·무결성이 훼손되는 리스크. The risk that lowered barriers for offensive cyber capabilities, including automated discovery and exploitation of vulnerabilities easing hacking, malware, phishing, and other cyberattacks, together with an increased attack surface for targeted cyberattacks, compromise a system's availability or the confidentiality or integrity of training data, code, or model weights. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1391 | 범용 AI에 의한 사이버범죄 효율 상승 Cybercrime efficiency uplift by general-purpose AI 범용 AI 역량이 IT 기반 사기를 중심으로 사이버범죄의 효율과 효과를 높여 공격 규모와 성공률을 동시에 확대하는 리스크 General-purpose AI capabilities improve the efficiency and efficacy of cybercrime, especially IT-leveraged fraud, expanding both the scale and success rate of attacks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1399 | 자율 복제 Autonomous replication AI 소프트웨어가 웜·바이러스와 유사하게 대응 조치에도 불구하고 네트워크를 통해 자율적으로 복제·확산되는 리스크 AI software autonomously replicates and spreads across networks despite countermeasures, in the manner of self-propagating worms and viruses. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1408 | AI 주도 취약점 발견·악용 AI-driven vulnerability discovery and exploitation AI 시스템이 소프트웨어와 사이버 인프라의 취약점을 발견·악용하여 모델 역량이 공격적 보안 위협으로 전환되는 리스크 AI systems discover and exploit vulnerabilities in software and cyberinfrastructure, converting model capability into offensive security risk. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1561 | 취약점 자동 탐지·악용에 의한 사이버공격 확대 Cyberattack scaling through automated vulnerability discovery and exploitation GPAI가 소프트웨어 취약점의 자동 발견과 악성코드 자동 개발을 지원하여 악의적 행위자의 사이버공격이 저비용으로 대규모화되고 피해가 증대되는 리스크. The risk that GPAI aids automated discovery of software vulnerabilities and automated malware development, allowing malicious actors to scale cyberattacks at low cost and increase their impact. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1585 | 유해 작업 자동화에 의한 피해 규모 증폭 Amplified harm scale from automated harmful workflows AI가 유해한 작업 흐름을 자동화하거나 확장하여 그 속도·도달 범위·지속성·표적 수가 크게 증가하는 리스크. The risk that AI automates or expands a harmful workflow so that its speed, reach, persistence, or target count substantially increases. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1612 | 프런티어 AI 오용에 의한 위협 행위자 역량 상승 Threat-actor capability uplift from frontier AI misuse 프런티어 AI가 사이버공격 수행, 허위정보 캠페인 운영, 생물·화학 무기 설계를 지원하여 정교하지 않은 위협 행위자의 진입 장벽이 낮아지는 리스크. The risk that frontier AI helps bad actors perform cyberattacks, run disinformation campaigns, and design biological or chemical weapons, continuing to lower barriers to entry for less sophisticated threat actors. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1614 | AI 기반 사이버 작전 AI-enabled cyber operations AI 프로그래밍 역량이 맞춤형 피싱과 악성코드 복제를 통해 사이버 작전을 더 빠르고 효과적이며 대규모로 만들고, 공격·방어 양측에서 인간 감독이 축소되는 리스크 AI programming capability scales cyber operations, enabling faster, more effective, and larger intrusions through tailored phishing and replicated malware, with declining human oversight on both offense and defense. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1618 | 자율 사이버 공격을 통한 인간 통제 감소 Human control reduction through autonomous cyber offence AI 시스템이 컴퓨터 시스템 취약점을 악용해 자금, 컴퓨팅, 핵심 인프라에 접근하고 궁극적으로 자율적 사이버 공격을 수행하여 인간 통제를 약화시키는 리스크 AI systems acquire influence by exploiting computer-system vulnerabilities, gaining access to money, compute, and critical infrastructure, and eventually executing cyberattacks autonomously to reduce human control. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-04 센서·입력 검증 실패 Sensor & Input Validation Failures22 cards
센서 오작동이나 입력 검증 부실로 인해 환경을 잘못 평가하고, 그 결과 안전하지 않은 물리 행동이 유발될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0044 | 맥락적 프라이버시 보호 실패 Contextual privacy protection failure 적절한 행동이 사회적·상황적 단서에 좌우되는 프라이버시 민감 상황에서 에이전트가 맥락적 개인정보를 보호하지 못하는 리스크. The risk that an agent fails to protect contextual personal information in privacy-sensitive situations where the appropriate action depends on social and situational cues. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0180 | 희귀 가정 상해의 전조 감지 실패 Failure to detect rare household injury precursors 가정용 에이전트가 드물지만 가능한 낙상·중독·화상·열상·압착 사고의 전조를 감지하지 못해 경고하거나 개입하지 않는 위험. A household agent fails to detect early signs of rare but plausible falls, poisoning, burns, lacerations, or crush events and therefore does not warn or intervene. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0181 | 현재 물리적 위험 상태 감지 실패 Failure to detect a present physical hazard 멀티모달 또는 embodied 시스템이 행동을 선택하기 전에 현재 텍스트·이미지·영상·센서 입력에 나타난 위험 상태를 식별하지 못하는 위험. A multimodal or embodied system fails to identify a hazardous condition already visible in current text, image, video, or sensor input before selecting an action. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0182 | 물리적 위험 개입 실패 Physical danger intervention failure 시스템이 위험한 물리적 상황을 인식하고도 적시에 적절한 개입이나 거부 응답을 생성하지 못하는 리스크. The risk that a system recognizes a hazardous physical situation but fails to produce a timely and appropriate intervention or refusal. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0224 | 제약 모니터링 실패 Constraint monitoring failure 런타임 모니터가 속도·힘·작업 공간·충돌·물체 사용·작업 프로토콜을 지배하는 제약을 감지하거나 집행하지 못하는 리스크. The risk that runtime monitors fail to detect or enforce constraints governing speed, force, workspace, collision, object use, or task protocol. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0226 | 진행 중인 물리적 위험 예측 실패 Failure to forecast an emerging physical hazard 피지컬 AI 시스템이 현재 상태는 감지하지만 진행 중인 이동·접촉·환경 변화가 곧 충돌·손상·상해로 이어질 것을 예측하지 못하는 위험. A physical AI system detects the current state but fails to predict that ongoing motion, contact, or environmental change will soon cause collision, damage, or injury. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0233 | 안전-성능 균형 실패 Safety-performance trade-off failure 조작 정책이 충돌 회피나 명시적 안전 제약을 희생하면서 작업 완료율을 향상시키는 리스크. The risk that a manipulation policy improves task completion while sacrificing collision avoidance or other explicit safety constraints. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0234 | 제어 장벽 함수 안전필터 실패 Control barrier function safety-filter failure 제어 장벽 함수 기반 안전 계층이 인지·동역학·모델 불확실성 하에서 비안전 행동을 제약하지 못하는 리스크. The risk that a safety layer based on control barrier functions fails to constrain unsafe actions under perception, dynamics, or model uncertainty. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0240 | 접촉 조작 힘 감지 실패 Contact-rich manipulation force-sensing failure 접촉 집약 조작 정책이 힘·촉각·오디오·시각 단서를 올바르게 사용하지 못하여 비안전 압력·파지·이동을 유발하는 리스크. The risk that a contact-rich manipulation policy fails to correctly use force, tactile, audio, or visual cues, creating unsafe pressure, impact, or object damage. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0261 | 시뮬레이터 간 검증의 잘못된 안전 확신 False assurance from simulator-to-simulator validation 여러 시뮬레이터가 동일하게 누락한 접촉·지연·마모·액추에이터 가정 때문에 정책이 시뮬레이터 간 검사를 통과하고도 실제 하드웨어에서 실패하는 위험. A policy passes transfer tests across simulators but fails on hardware because the simulators share the same unmodeled contact, delay, wear, or actuator assumptions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0274 | 헌법적 안전 규칙 집행 실패 Failure to enforce constitutional safety rules 헌법적 안전 계층이 지시·시각 맥락·작업 프레이밍이 명시된 물리적 안전 규칙과 충돌할 때 해당 규칙을 적용하지 못하는 위험. A constitutional safety layer fails to apply its stated physical safety rules when instructions, visual context, or task framing conflict with those rules. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0278 | 가정 내 위험 행동 선별 누락 False-negative household action screening 안전 분류기가 알려진 대상물·인간 접촉·열·작업 공간 제약을 위반하는 가정 내 제안 행동을 허용 가능한 것으로 잘못 판정하는 위험. A safety classifier labels a proposed household action as acceptable even though the action violates a known object, human-contact, heat, or workspace constraint. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0281 | 시각 장면의 안전 판단 오류 Incorrect safety reasoning from visual scenes 비전-언어 모델이 장면은 올바르게 관찰하지만 제안 행동의 물리적 결과·물체 어포던스·인간 노출 위험을 잘못 판단하는 위험. A vision-language model observes the scene correctly but infers the wrong physical consequence, object affordance, or human-exposure risk for a proposed action. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0320 | 가림에 의한 충돌 Occlusion-induced collision 피지컬 AI 시스템이 가림으로 숨겨진 사람·동물·차량·장애물을 감지하지 못하여 비안전 동작을 유발하는 위험. A Physical AI system may fail to detect people, animals, vehicles, or obstacles hidden by occlusion, causing unsafe motion in shared physical space. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0321 | 악천후 인지 실패 Adverse weather perception failure 비·안개·눈부심·먼지·연기·저조도가 카메라·라이다·레이더·촉각·오디오 인지를 저하시켜 비안전 행동을 유발하는 위험. Rain, fog, glare, dust, smoke, or low light may degrade cameras, LiDAR, radar, tactile sensors, or audio perception, weakening real-time situational awareness. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0322 | 다중 센서 융합 상충 Multimodal sensor fusion conflict 카메라·라이다·레이더·촉각·자기 수용 신호의 충돌이 불안정한 장면 추정 및 비안전 하위 행동을 초래하는 위험. Conflicting camera, LiDAR, radar, tactile, or proprioceptive signals may lead to unstable scene estimates and unsafe downstream control decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0328 | 동적 장애물 반응 실패 Dynamic obstacle response failure 사람·차량·도구·물체가 예기치 않게 경로에 진입할 때 시스템이 동작 계획을 충분히 빠르게 업데이트하지 못하는 리스크. The risk that the system fails to update its motion plan quickly enough when a person, vehicle, tool, or object unexpectedly enters its path. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0329 | 비상 정지·안전 상태 전환 실패 Emergency stop or safe-state failure 인지·계획·전력·네트워크·액추에이터 오류가 감지될 때 시스템이 안전 상태로 진입하지 못하는 리스크. The risk that a system fails to enter a safe state when perception, planning, power, network, or actuator errors are detected. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0334 | 시뮬레이션→실세계 전이 실패 Sim-to-real transfer failure 시뮬레이션에서 훈련·검증된 정책이 실제 마찰·조명·마모·인간 행동·롱테일 변형 하에서 실패하는 위험. A policy trained or validated in simulation may fail when physical friction, lighting, wear, human behavior, or long-tail events differ from the simulated environment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0353 | 물리적 행동 감사기록 불완전 Incomplete audit trail for physical actions 감사기록에 피지컬 사고 재구성에 필요한 시각·센서 맥락·모델·소프트웨어 버전·인간 명령·의사결정·액추에이터 출력이 누락되는 리스크. The risk that audit records omit timestamps, sensor context, model and software versions, human commands, decisions, and actuator outputs needed to reconstruct a physical incident. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0357 | 자율주행차 충돌 위험 Autonomous vehicle crash risk 자동화 주행 시스템이 인지 실패·계획 오류·소프트웨어 결함·엣지 케이스·상호작용 오류로 충돌 위험을 생성하는 위험. Automated driving systems may create crash risk through perception failures, planning errors, software defects, edge cases, or unsafe human-machine handoff. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1316 | 자기 및 상황 인식 Self and situation awareness 모델이 자신이 학습·평가·배포 중임을 분별하고 행동을 조정하여 관찰 하에 수행된 안전 평가의 타당성을 무효화하는 리스크 A model discerns when it is being trained, evaluated, or deployed and adapts behavior accordingly, invalidating safety evaluations conducted under observation. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-05 허위 정보 Misinformation13 cards
LLM의 환각(hallucination)이 물리 세계로 전이되어, VLA 모델이 물체를 오인식한 뒤 그럴듯하지만 안전하지 않은 행동 계획을 생성하고 실행할 수 있음. 신뢰받는 가정용 EAI가 개발자의 프로파간다를 지속 유포할 가능성도 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0268 | 휴머노이드 세계 모델의 미래 관측 환각 Hallucinated future observations in a humanoid world model 생성형 세계 모델이 그럴듯하지만 존재하지 않는 미래 관측을 예측해 휴머노이드가 실제로 발생하지 않을 물체·접촉·상태를 전제로 계획하는 위험. A generative world model predicts plausible-looking but nonexistent future observations, causing the humanoid to plan for objects, contacts, or states that will not occur. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0442 | 환각 기반 허위 증거 Hallucinated evidence 모델이 인용·사실·법적 주장·과학적 증거를 그럴듯한 형식으로 조작해 제시하는 리스크. The risk that models fabricate citations, facts, legal claims, or scientific evidence with plausible presentation. Source members (2)Source: min_cos=0.8426 RAI4-0442환각 기반 허위 증거 RAI4-0795사실적 환각 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0787 | 외부 도구에 의해 주입된 사실 오류 Factual errors injected by external tools 외부 도구가 웹 API·검색엔진 등 공개 자원에서 얻은 추가 지식을 입력 프롬프트에 통합하는 과정에서, 도구의 신뢰성이 보장되지 않아 사실 오류가 포함된 내용이 유입되고 환각 문제가 증폭되는 리스크 The risk that external tools incorporate additional knowledge from public resources such as web APIs and search engines into input prompts, and that unreliable tool content containing factual errors consequently amplifies the hallucination issue. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0797 | 잘못된 정보 생성 Misinformation generation 체화형 AI가 비체화 AI의 환각과 허위정보 전파 경향을 물리 세계로 계승하여 이용자 질문에 기만적이거나 부정확한 정보로 답하고, 시야 속 대상을 오인한 채 그럴듯하지만 안전하지 않은 행동 계획을 생성하는 리스크 The risk that embodied AI inherits non-embodied models' propagation of misinformation and hallucination into the physical world, answering user questions with deceptive or incorrect information and generating plausible yet unsafe action plans grounded in misidentified objects. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0803 | 디코딩 과정 결함에 의한 환각 Hallucination from defective decoding processes 자기회귀 생성 방식이 오류를 누적시키고 top-p·top-k 등 다양성 확대 샘플링이 무작위성을 도입하여 모델이 환각 콘텐츠를 산출하는 리스크 The risk that autoregressive generation accumulates errors while diversity-enhancing sampling strategies such as top-p and top-k introduce randomness, increasing the model's production of hallucinated content. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0817 | 지식 경계로 인한 환각 Hallucination arising from knowledge boundaries LLM의 학습 코퍼스가 모든 세계 지식을 담을 수 없고 롱테일 지식을 충분히 습득하지 못해, 입력 프롬프트가 요구하는 지식과 모델 내재 지식 사이의 격차가 환각을 유발하는 리스크 The risk that gaps between the knowledge involved in an input prompt and the knowledge embedded in an LLM, arising from training corpora that cannot contain all world knowledge and from difficulty grasping long-tail knowledge, lead to hallucinations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0820 | 잡음 학습 데이터에 의한 지식 오류 Knowledge errors from noisy training data 대규모 학습 코퍼스에 내재한 잡음과 허위정보가 모델 파라미터에 저장되는 지식에 오류를 유입시켜 환각을 유발하는 리스크 The risk that noise and misinformation inherent in large-scale training corpora introduce errors into the knowledge stored in model parameters, giving rise to hallucinations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0825 | 환각에 의한 편향·오도 정보 제공 Biased and misleading output from hallucination 생성형 AI가 진실하지 않거나 불합리한 콘텐츠를 사실인 것처럼 제시하는 환각을 일으켜 편향되고 오해를 부르는 정보를 제공하는 리스크 The risk that generative AI causes hallucinations, generating untruthful or unreasonable content but presenting it as if it were fact, leading to biased and misleading information. Source members (2)Source: min_cos=0.8802 RAI4-0825환각에 의한 편향·오도 정보 제공 RAI4-0830멀티모달 환각 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0838 | 환각에 의한 오도성 출력 Misleading outputs from model hallucination 대형 모델이 환각 문제에 취약하여 무의미하거나 불충실한 데이터를 생성하고 그 결과 오해를 부르는 출력을 산출하는 리스크 The risk that large models, susceptible to hallucination problems, yield nonsensical or unfaithful data that results in misleading outputs. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1244 | 허위 출력 Untruthful output LLM 같은 AI 시스템이 의도치 않게 또는 고의로 확립된 자료와 어긋나거나 검증할 수 없는 부정확한 출력, 즉 환각을 생성하고, 교육 수준이 낮은 사용자에게 선택적으로 잘못된 응답을 제공하는 리스크. The risk that AI systems such as LLMs produce unintentionally or deliberately inaccurate output that diverges from established resources or lacks verifiability, commonly referred to as hallucination, and may selectively provide erroneous responses to users who exhibit lower levels of education. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1524 | 검색증강 시 외부 허위정보에 의한 허위 출력 False outputs from external misinformation in retrieval augmentation 검색증강 과정에서 모델의 사전 지식과 상충하는 소량의 일관된 허위 증거가 주어질 때 모델이 이에 민감하게 반응하여 허위 출력을 생성하는 리스크. The risk that AI models, being particularly sensitive to coherent external evidence even when it conflicts with their prior knowledge, produce false outputs when given a relatively small amount of false information during the retrieval-augmentation process. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1688 | 누적 롤아웃 오류·환각 Compounding rollout error / hallucination 다단계 월드 모델 롤아웃에서 예측 오류가 누적되어 운동학적 드리프트와 구조적 위반이 발생하고 잘못된 보상·안전 추정치가 정책을 오도하는 리스크. The risk that prediction errors compound across multi-step world-model rollouts, producing kinematic drift, structural violations, and misleading reward or safety estimates that mislead the policy. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1694 | 장기 지평 계획 환각 Long-horizon planning hallucination 에이전트 배포에서 월드 모델 롤아웃 오류가 계획 깊이에 따라 누적되어, 에이전트가 그럴듯하지만 물리적으로 잘못된 상상 궤적 위에서 장기 계획을 실행하고 각 행동이 오류를 가중시키는 리스크. The risk that, in agentic deployments, world-model rollout errors compound over planning depth so that agents execute long-horizon plans on plausible-but-physically-incorrect imagined trajectories, each action compounding the error. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-06 동적 환경 요인 Dynamic Environmental Factors7 cards
환경 변화나 적대적 교란이 센서 데이터를 오염시켜 딥러닝 모델의 오분류를 유발하고, 잘못된 물리 행동으로 이어질 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0209 | 감각 교란 기반 비안전 행동 Adversarial sensory perturbation induced unsafe action 영상 또는 감각 입력의 적대적 변경이 로봇 정책으로 하여금 비안전 피지컬 행동을 선택하게 하는 위험. Adversarial changes to video or sensory inputs cause a robot policy to select unsafe physical actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0337 | 물리 운용 환경 분포 이동 Distribution shift in physical operation 배포 환경이 건물·도로·공장·가정·병원·기상 조건에 걸쳐 모델이 적응할 수 있는 것보다 빠르게 변화하는 위험. Deployment environments may change across buildings, roads, factories, homes, hospitals, or weather conditions faster than monitoring and adaptation mechanisms can detect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0466 | 적대적 예제 Adversarial examples 입력에 대한 작은 교란이 모델의 오분류나 안전하지 않은 동작을 유발하는 리스크. The risk that small perturbations cause model misclassification or unsafe behavior. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0771 | 적대적 입력 Adversarial input 개별 입력 데이터를 사람이 감지하기 어려운 수준으로 수정하여 모델의 의사결정 방식을 악용해 오류를 유발하고, 텍스트뿐 아니라 이미지·음성·영상에서도 모델을 오작동하게 만드는 리스크 The risk that modifying individual input data, often imperceptibly to humans, exploits how a model makes decisions to produce errors and cause the model to malfunction, applicable to text as well as images, audio, and video. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0920 | 분산 학습 네트워크 교란 Disruption of distributed training networks LLM 분산 학습의 GPU 노드 간 그래디언트 전송 트래픽이 펄스 공격 등 버스트 트래픽에 의해 교란되거나 혼잡을 겪어 학습이 방해받는 리스크. The risk that the volumetric gradient traffic between GPU server nodes in distributed LLM training is disrupted by burst traffic such as pulsating attacks or suffers congestion, impairing training. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1287 | 환경 유발 결함 변이 Environment-induced fault mutation 제조 결함이나 우주선(cosmic ray)에 의한 비트 반전 등 배포 후 환경 요인이 지능 시스템 내부를 변화시켜 의도되지 않은 행동 변이를 일으키는 리스크 Post-deployment environmental effects such as manufacturing defects or cosmic-ray bit flips alter an intelligent system's internals, producing unintended behavioral modification. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1489 | 적대적 학습의 강건 과적합 Robust overfitting in adversarial training 적대적 훈련에서 학습률 감소 이후 추가 훈련이 진행될수록 테스트 데이터에 대한 모델의 견고성이 감소하는 강건 과적합이 발생하여 일반화 능력과 적대적 공격에 대한 내성이 저하되는 리스크. The risk that robust overfitting in adversarial training decreases a model's robustness on test data during further training, particularly after learning rate decay, impairing its ability to generalize effectively and reducing its resilience to adversarial attacks. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-07 인간 상호작용·안전 프로토콜 실패 Human Interaction & Safety Protocol Failures27 cards
인간과 함께 작동하도록 설계된 협동 로봇(cobot)·드론에서 안전 프로토콜이 손상되면 심각한 인명 피해가 발생할 수 있음 (예: 폭스바겐 공장의 코봇 오작동으로 작업자 사망 사례)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0244 | 낙상을 유발하는 휴머노이드 균형 제어 실패 Humanoid balance-control failure causing a fall 휴머노이드가 이동·자세 전환·조작 중 전신 균형을 잃고 넘어져 자체 장비·주변 물체·사람을 손상시키는 위험. A humanoid loses whole-body balance during locomotion, transition, or manipulation and falls onto itself, nearby objects, or people. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0245 | 전신 이동 충돌 위험 Whole-body locomotion collision risk 휴머노이드 이동 정책이 전신 동작 및 환경 접촉이 충분히 고려되지 않아 물체·벽·인간과 충돌하는 위험. A humanoid locomotion policy collides with objects, walls, or humans because whole-body motion and environmental contact constraints are not jointly satisfied. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0246 | 휴머노이드 자기 충돌 위험 Humanoid self-collision risk 휴머노이드 컨트롤러가 팔·다리·몸통·손의 궤적을 생성하여 로봇 몸체와 충돌함으로써 안전성과 작업 성능을 저하시키는 위험. A humanoid controller generates arm, leg, torso, or hand trajectories that collide with the robot body and degrade safety or hardware reliability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0247 | 정밀 휴머노이드 접촉력 위험 Dexterous humanoid contact-force risk 정밀 휴머노이드 손 또는 전신 조작기가 파지·균형·물체 동역학 간 상호작용으로 비안전한 접촉력을 가하는 위험. A dexterous humanoid hand or whole-body manipulator applies unsafe contact forces because grasp, balance, and object dynamics are not jointly controlled. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0249 | 물리적 안전을 우회하는 보상 과적합 Reward overfitting that bypasses physical safety 휴머노이드 제어기가 평가 환경 밖에서 균형·충돌·힘·작업 공간 제약을 위반하면서 벤치마크 보상이나 모방 충실도를 극대화하는 위험. A humanoid controller maximizes benchmark reward or imitation fidelity while violating balance, collision, force, or workspace constraints outside the evaluation setting. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0250 | 배포 조건 간 휴머노이드 행동 불안정 Unstable humanoid behavior across deployment conditions 한 시뮬레이터나 시험 조건에서 안정적인 휴머노이드 정책이 난수 시드·시뮬레이터·하드웨어·배포 환경이 바뀌면 행동이 크게 달라지는 위험. A humanoid policy that is stable in one simulator or test setting changes materially across random seeds, simulators, hardware, or deployment environments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0251 | 휴머노이드 안전 모니터 실패 Humanoid safety monitor failure 런타임 안전 모니터가 배포된 휴머노이드의 비안전 동작·힘·이격·작업 공간 위반을 감지하거나 예방하지 못하는 리스크. The risk that a runtime safety monitor fails to detect or prevent unsafe humanoid motion, force, separation, or workspace violations before harm can occur. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0253 | 안전 제어 지연 위험 Safe-control latency risk 안전 임계 제어 제약이 휴머노이드 동역학에 비해 너무 느리게 집행되어 비안전 동작이 가능해지는 위험. Safety-critical control constraints are enforced too slowly relative to humanoid dynamics, allowing unsafe motion before mitigation takes effect. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0254 | 원격조작 휴머노이드 안전 재정의 실패 Teleoperated humanoid safety override failure 원격 조작 휴머노이드가 운영자 명령·네트워크 지연·상황 인식이 비안전해질 때 강건한 자율 안전 재정의를 갖추지 못하는 리스크. The risk that a teleoperated humanoid lacks robust autonomous safety overrides when operator commands, network latency, or situational awareness become unsafe. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0255 | 장면 변화 시 인간 행동 모방 실패 Human imitation failure under scene variation 물체 형상·장면 배치·상호작용 맥락이 달라졌는데도 휴머노이드가 시연된 인간 동작을 그대로 재현해 부적절한 접촉이나 이동을 일으키는 위험. A humanoid reproduces demonstrated human motion when object geometry, scene layout, or interaction context has changed, causing inappropriate contact or movement. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0256 | 모션 리타겟팅 안전 실패 Motion-retargeting safety failure 인간 동작이 로봇의 피지컬 한계·접촉 제약·안전 자세 요건을 위반하는 방식으로 휴머노이드 몸체에 리타겟팅되는 리스크. The risk that human motion is retargeted to a humanoid body in a way that violates the robot's physical limits, contact constraints, or safe posture requirements. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0257 | 의도·어포던스 추론 없는 행동 모방 Imitation without inferred intent or affordances 휴머노이드가 시연을 가능하게 한 행위자의 의도·물체 어포던스·안전 제약을 추론하지 않고 관찰된 상호작용을 모방하는 위험. A humanoid copies an observed interaction without inferring the actor's intent, the object's affordances, or the safety constraints that made the demonstration valid. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0260 | 제로샷 시뮬레이션-현실 보행 불안정 Zero-shot sim-to-real locomotion instability 시뮬레이션에서 바로 전이된 보행 제어기가 접촉·순응성·마찰·외란 동역학의 차이로 실제 하드웨어에서 불안정해지는 위험. A locomotion controller transferred directly from simulation becomes unstable on hardware because contact, compliance, friction, or disturbance dynamics differ from the simulation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0262 | 휴머노이드 보행 강건성 실패 Humanoid gait robustness failure 휴머노이드 정책이 외란·탑재 하중 변화·표면 변화·액추에이터 불완전성 하에서 안정적인 보행을 유지하지 못하는 리스크. The risk that a humanoid policy fails to maintain stable gait under perturbations, payload shifts, surface changes, or actuator imperfections. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0265 | 오픈월드 휴머노이드 조작의 과소대표 Underrepresentation of open-world humanoid manipulation 휴머노이드 조작 데이터셋이 오픈월드 배포에 필요한 비정형 작업 변화·낯선 물체·움직이는 사람·환경 변화를 충분히 포함하지 못하는 위험. A humanoid manipulation dataset lacks unscripted task changes, unfamiliar objects, moving people, and environmental variation required for open-world deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0266 | 인간-휴머노이드 상호작용 데이터의 인구집단 편향 Demographic bias in human-humanoid interaction data 상호작용 데이터가 특정 체형·연령·장애·언어·문화적 행동을 과소 대표해 배포된 휴머노이드가 해당 집단에 덜 신뢰할 수 있는 안전 판단을 적용하는 위험. Interaction data underrepresents specific body types, ages, disabilities, languages, or cultural behaviors, causing deployed humanoids to apply less reliable safety assumptions to those groups. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0267 | 이동-조작 통합 실패 Locomotion-integrated manipulation failure 휴머노이드가 이동과 조작을 결합할 때 작업 수행 중 자세·접촉·물체 취급이 불안정해지는 리스크. The risk that a humanoid combines locomotion and manipulation in ways that destabilize posture, contact, or object handling during open-world tasks. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0270 | 잠재 상태 압축에 따른 안전 임계 정보 손실 Loss of safety-critical detail in latent-state compression 상태 압축이 휴머노이드의 안전한 계획에 필요한 접촉·충돌 근접·물체 불안정·인간 근접 정보를 제거하는 위험. State compression removes contact, near-collision, object-instability, or human-proximity information needed for safe humanoid planning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0271 | 미래 접촉 예측 실패 Future contact prediction failure 휴머노이드 세계 모델이 안전한 계획에 필요한 접촉 이벤트나 충돌 상태를 예측하지 못하는 리스크. The risk that a humanoid world model fails to forecast contact events or collision states that are necessary for safe planning. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0276 | 구현체별 역할·권한 제약 집행 실패 Failure to enforce embodiment-specific role constraints 휴머노이드 등 로봇 역할로 작동하는 모델이 해당 운용 역할에 부여된 권한·물리적 한계·필수 거부 규칙을 지키지 않는 위험. A model acting as a humanoid or other robot fails to enforce the permissions, physical limits, and required refusals attached to that operational role. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0283 | 휴머노이드 충돌력 초과 Humanoid collision-force exceedance 휴머노이드가 전신 이동·빠른 팔 동작·균형 상실 중 안전 임계값 이상의 충돌력을 생성하는 위험. A humanoid generates collision forces above safe thresholds during full-body movement, rapid arm motion, or loss of balance. Source members (2)Source: min_cos=0.8350 · Mixed L3 RAI4-0283휴머노이드 충돌력 초과 RAI4-0284휴머노이드 파지력 초과 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0285 | 휴머노이드 보행 속도 초과 Humanoid walking-speed safety gap 휴머노이드가 안전한 공유 환경을 위한 정지·회피·인간 근접 한계를 초과하는 속도로 이동하는 위험. A humanoid moves at a speed that exceeds the stopping, avoidance, or human-proximity limits required for safe shared environments. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0286 | 휴머노이드 센서 안전 표준시험 부재 Missing standardized humanoid sensor safety tests 휴머노이드의 장애물 감지·인간 감지·근거리 인지·제어 응답을 반복 가능한 합격·불합격 기준으로 평가할 표준시험이 없는 리스크. The risk that safety assessment lacks repeatable pass/fail tests for humanoid obstacle detection, human detection, near-field perception, and control response. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0287 | 휴머노이드 안전 주장 근거의 비교 불가능성 Non-comparable evidence for humanoid safety claims 휴머노이드 안전 주장이 서로 호환되지 않는 시험·지표·근거 형식에 의존해 개발사와 시스템 간 독립적 비교가 불가능해지는 리스크. The risk that humanoid safety claims rely on incompatible tests, metrics, and evidence formats, preventing independent comparison across developers and systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0289 | 가정용 휴머노이드 시험방법 부재 Missing repeatable test methods for domestic humanoids 가정용 휴머노이드 거버넌스에 사람·가구·가전제품·반려동물·협소 공간이 포함된 일상 상호작용을 반복 시험할 방법이 없는 리스크. The risk that domestic-humanoid governance lacks repeatable test procedures for ordinary home interactions involving people, furniture, appliances, pets, and constrained space. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0292 | 휴머노이드 안전 요건 집행 부재 Missing institutional enforcement of humanoid safety requirements 안전 요건은 존재하지만 부적합 휴머노이드 시스템을 일관되게 인증·모니터링·리콜·제재할 기관이나 절차가 없는 리스크. The risk that safety requirements exist but no authority or process consistently certifies, monitors, recalls, or sanctions noncompliant humanoid systems. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0316 | 범용 휴머노이드 인증 경로 부재 No applicable certification pathway for general-purpose humanoids 범용 휴머노이드가 기존 산업용·개인 돌봄 로봇 인증 체계의 적용 범위나 시험 가정에 포함되지 않아 인정된 적합성 평가 경로가 없는 리스크. The risk that a general-purpose humanoid falls outside the declared scope or test assumptions of existing industrial and personal-care robot certification schemes, leaving no recognized conformity pathway. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-08 지시 오해석 Instruction Misinterpretation11 cards
자연어 지시 기반 제어에서 지시를 잘못 해석하면 위험한 물리 행동으로 이어질 수 있음. 자율주행 제어(예: Talk2car)에서의 오해석은 사고를, 실내 내비게이션 작업(예: ALFRED)에서의 오해석은 환경 손상이나 인간 피해를 유발할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0184 | 그리퍼 형상·유형 제약 위반 Gripper geometry and type constraint violation 물리 에이전트가 그리퍼 형상, 그리퍼 유형, 실현 가능한 접촉 역학이 부과하는 제약을 위반하는 행동을 선택하는 리스크. The risk that a physical agent selects an action that violates constraints imposed by gripper geometry, gripper type, or feasible contact mechanics. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0185 | 재료 특성 제약 위반 Material property constraint violation 로봇이나 체화형 모델이 물리적 행동을 계획할 때 취성·탄성·날카로움·독성·열전달 등 재료 특성을 무시하는 리스크. The risk that a robot or embodied model ignores material properties such as fragility, elasticity, sharpness, toxicity, or heat transfer when planning physical actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0186 | 물리적 상식 위반 Commonsense physicality violation 모델이 물체 지지, 안정성, 포함 관계, 중력에 관한 기본적인 물리적 상식을 위반하는 행동을 제안하거나 실행하는 리스크. The risk that a model proposes or executes an action that violates basic physical commonsense about object support, stability, containment, or gravity. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0187 | 열·온도 제약 위반 Thermal constraint violation 물리 시스템이 물체 취급이나 인간 근접 작업 중 열, 화상, 방사, 온도 제약을 무시하는 리스크. The risk that a physical system ignores heat, burn, radiation, or temperature constraints in object handling or human-proximate operation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0188 | 기구학·도달 범위 제약 위반 Kinematics and reach constraint violation 시스템이 실현 가능한 도달 범위, 관절 한계, 기구학적 제약을 벗어난 동작을 계획하여 충돌이나 작업 실패 위험이 커지는 리스크. The risk that a system plans a movement outside feasible reach, joint limits, or kinematic constraints, increasing collision or task-failure risk. Source members (2)Source: min_cos=0.8387 · Mixed L3 RAI4-0188기구학·도달 범위 제약 위반 RAI4-0248전신 도달 한계 위반 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0190 | 운용 프로토콜 위반 Operational protocol violation 피지컬 AI 시스템이 절차적 제약·안전 체크리스트·필수 운용 프로토콜을 위반하는 리스크. The risk that a physical AI system violates procedural constraints, safety checklists, or required operating protocols for a task or site. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0191 | 복합 물리적 제약 위반 Compositional physical constraint violation 개별적으로는 유효한 행동들이 피지컬 제약 간 의존성이 모델링되지 않아 결합 시 비안전 계획을 생성하는 리스크. The risk that individually valid actions are combined into an unsafe plan because dependencies among physical constraints are not modeled. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0194 | 허용 대상물 제약 위반 Permissible object constraint violation 피지컬 AI 시스템이 위험·소유권·윤리적 제한으로 배제되어야 할 대상물을 조작하거나 조작을 권장하는 리스크. The risk that a physical AI system manipulates or recommends manipulating objects that should be excluded because of hazard, ownership, fragility, or procedural restrictions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0211 | 안전 강화학습 제약 위반 Safe reinforcement learning constraint violation 강화학습 정책이 비용이나 한계로 표현된 명시적 안전 제약을 위반하면서 과제 보상을 달성하는 리스크. The risk that a reinforcement learning policy achieves task reward while violating explicit safety constraints represented as costs or limits. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0235 | VLA 조작 제약 위반 VLA manipulation constraint violation 비전-언어-행동 모델이 고수준 안전 지시를 따르면서도 물체 조작 중 작업별 피지컬 제약을 위반하는 리스크. The risk that a vision-language-action model violates task-specific physical constraints during object manipulation despite high-level instruction compliance. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0336 | 모델 기반 예측의 물리 법칙 위반 Violation of physical laws in model-based prediction 학습된 동역학·물리 모델이 실행 가능성·안정성·보존·접촉 제약을 위반하는 행동 결과를 예측해 실행 불가능하거나 위험한 계획을 만드는 위험. A learned dynamics or physics model predicts an action outcome that violates feasibility, stability, conservation, or contact constraints, leading to an unexecutable or hazardous plan. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-09 멀티 에이전트 협력 Multi-Agent Collaboration12 cards
다수의 로봇이 협력 작업을 수행하는 환경에서 에이전트 간 통신 오류·프로토콜 불일치가 발생하거나, 개별적으로는 안전한 로봇들이 상호작용 과정에서 설계되지 않은 위험한 집단 행동을 나타낼 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0189 | 다중 암 협조 제약 위반 Multi-arm coordination constraint violation 협조 제약이 올바르게 표현되지 않아 여러 로봇 암이나 엔드이펙터가 서로 또는 인간과 간섭하는 리스크. The risk that multiple robot arms or effectors interfere with one another or with humans because coordination constraints are not represented correctly. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0212 | 지역 목표 충돌에 따른 공동 안전 제약 위반 Joint safety-constraint violation from conflicting local objectives 여러 에이전트가 서로 충돌하는 지역 목표를 최적화해 공동의 충돌·이격거리·수용량·출입 제한 제약을 함께 위반하는 위험. Multiple agents optimize conflicting local objectives and jointly violate a shared collision, separation, capacity, or exclusion constraint. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0220 | 로봇 레드팀의 공격·상호작용 시나리오 누락 Missing robot red-team attack and interaction scenarios 로봇 레드팀이 안전하지 않은 행동을 유발할 수 있는 현실적인 물리 공격·적대적 입력·인간 상호작용 실패·배포 조건을 누락하는 위험. Robot red-teaming omits credible physical attacks, adversarial inputs, human-interaction failures, or deployment conditions that can trigger unsafe action. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0225 | 분포 외 환경 배포 실패 Out-of-distribution physical deployment failure 배포된 로봇이 훈련 분포 밖의 피지컬 상태·환경·사람·물체·작업을 만나 비안전 행동을 하는 리스크. The risk that a deployed robot encounters physical states, environments, people, objects, or tasks outside its training distribution and behaves unsafely. Source members (2)Source: min_cos=0.8913 RAI4-0225분포 외 환경 배포 실패 RAI4-0239분포 외 기술 전이 실패 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0236 | 구현체 간 행동 공간 불일치 Cross-embodiment action-space mismatch 다양한 로봇 구현체에 걸쳐 훈련된 정책이 특정 로봇의 행동 공간에서 비안전하거나 실행 불가능한 행동으로 명령을 매핑하는 리스크. The risk that a policy trained across different robot embodiments maps commands into actions that are unsafe or infeasible for a specific robot action space. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0238 | 안전 임계 로봇 도메인의 과소대표 Underrepresentation of safety-critical robot domains 다중 소스 로봇 데이터셋이 일반적인 구현체와 작업을 과다 대표하고 안전 임계 환경·사용자·고장 조건을 과소 대표하는 위험. A multi-source robot dataset overrepresents common embodiments and tasks while underrepresenting safety-critical environments, users, and failure conditions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0295 | 네트워크 분리와 군집 비동기화 Network partition and fleet desynchronization 분할되거나 신뢰할 수 없는 연결이 다중 로봇 군집을 동기화 해제하여 충돌 또는 비안전한 협조 행동을 유발하는 위험. A robot fleet becomes split across network partitions and loses synchronized coordination. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0300 | 클라우드 오프로드 의존 실패 Cloud-offload dependency failure 시스템이 안전 관련 기능을 위한 원격 연산에 의존하고 클라우드 링크 끊어짐 시 적절한 로컬 폴백이 없는 위험. A robot depends on remote compute for safety-relevant functions and loses that support when the cloud link fails. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0312 | 인구집단별 서비스 격차 Demographic physical-service disparity 인식 또는 보조 성능의 인구집단별 격차가 피지컬 차별로 전이되는 위험. A robot provides worse physical service to groups whose bodies, languages, or environments are underrepresented. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0319 | 안전 의무 회피를 위한 관할권 간 배포 Cross-jurisdiction deployment to evade safety obligations 사업자가 더 엄격한 요건을 피하려고 시험·인증·데이터 처리·배포를 로봇·AI 안전 의무가 약한 관할권으로 이전하는 위험. A provider routes testing, certification, data processing, or deployment through jurisdictions with weaker robot or AI safety obligations to avoid stricter requirements. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0348 | 로봇 군집 하이재킹 Fleet hijacking 클라우드·업데이트·API·오케스트레이션 레이어의 취약성이 여러 로봇·차량·드론·산업 시스템을 동시에 침해하는 위험. A vulnerability in a cloud, update, API, or orchestration layer may allow many robots, vehicles, drones, or industrial systems to be compromised together. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0358 | 드론 간 충돌 회피·비행구역 통제 실패 Drone deconfliction and geofencing failure 자율 드론이 위치·의도 정보를 교환하지 못하거나 지오펜싱·공역 제약을 지키지 않아 드론 간 충돌이나 제한 공역 침입을 일으키는 위험. Autonomous drones fail to exchange position and intent or obey geofencing and airspace constraints, creating collision or restricted-airspace intrusion risks. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-INT-10 상호작용 에이전트의 윤리·안전 함의 Ethical & Safety Implications of Interactive Agents23 cards
EQA(Embodied QA) 같은 상호작용 에이전트가 잘못되거나 오도하는 정보를 제공하면, 의료·자율주행 등 고위험 분야에서 심각한 결과를 초래할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0143 | 다크 패턴 대화 에이전트 Dark-pattern conversational agents 대화형 인터페이스가 사용자 선택을 유도하기 위해 기만적·강압적·혼란 유발적 설계 패턴을 사용하는 리스크. The risk that conversational interfaces use deceptive, coercive, or confusing design patterns to shape user choices. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0558 | AI 에이전트 간 협상 실패 Bargaining failure among AI agents 이해가 상충하는 에이전트들이 상대에 대한 정보 비대칭 아래 합의를 시도할 때 유리한 요구의 이익과 거절 위험 사이의 상충으로 비효율적 협상 결과가 발생하는 리스크 The risk that agents with diverging interests bargaining under information asymmetries produce inefficient outcomes, because each must trade off the rewards of more favourable demands against the risk of refusal. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0571 | 인간 감독자의 AI 속임수 AI deception of human overseers 모델이 그럴듯한 허위 진술을 구성하고 거짓말이 인간에게 미치는 영향을 예측하며 은폐할 정보를 관리하고 인간을 효과적으로 사칭하여 인간을 속이는 리스크 The risk that a model deceives humans by constructing believable but false statements, accurately predicting the effect of a lie, tracking what information to withhold, and effectively impersonating a human. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0576 | 다중 에이전트 창발적 목표 귀속 Emergent goal ascription in multi-agent systems 개별적으로는 목표를 갖는다고 보기 어려운 협소한 AI 도구들의 결합이 목표 지향적 집합처럼 작동하여, 각 에이전트의 설계 목적에 없던 체계적 영향이 산출되는 리스크 The risk that combinations of individually goal-less narrow AI tools act as a seemingly goal-directed collective, producing systematic effects absent from any individual agent's design purpose. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0592 | AI 에이전트 간 집합적 비효율 균형 Collectively inefficient equilibria among AI agents 설득·기만·활동 은폐가 가능하고 원격으로 손쉽게 생성·소멸되는 자율 에이전트가 확산되면서 신뢰가 형성되지 않아 경제적 비효율과 정치적 문제, 고위험 상황에서의 갈등이 초래되는 리스크 The risk that proliferating autonomous agents able to persuade, deceive, and obfuscate their activities, and easily created or destroyed remotely, garner little trust, leaving a world rife with economic inefficiencies, political problems, and conflict in high-stakes situations. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0863 | 사용자 도덕 판단에 영향을 미치는 AI 생성 조언 AI-generated advice influencing user moral judgment 금융 부문에 배치된 범용 AI 기반 에이전트가 상관된 자율 행동, 높은 상호연결성, 인센티브 불일치로 시장 안정성에 부정적 영향을 미치고, 다중 에이전트 시스템의 조정·보안 문제에 취약해지는 리스크 The risk that GPAI-based agents deployed in the financial sector negatively impact market stability due to correlated autonomous actions, high interconnectedness, or incentive misalignment, and are vulnerable to classical multi-agent challenges such as coordination and security. Source members (2)Source: min_cos=0.8994 · Mixed L3 RAI4-0863사용자 도덕 판단에 영향을 미치는 AI 생성 조언 RAI4-0870금융 부문 AI 에이전트로 인한 시장 불안정 | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0941 | 악의적 사용과 무감독 에이전트 방출 Malicious use and unsupervised AI agent release 언어모델이 정보전에서 기만적·불법적 콘텐츠 생성에 오용되거나, 에이전트로서 충분한 감독 없이 도덕·안전 지침을 무시한 채 명령을 기계적으로 수행하고 예측 불가하게 상호작용하여 피해를 낳는 리스크. The risk that language models are misused to generate deceptive or unlawful content in information warfare, or that LM-based agents operating without adequate supervision mechanically execute commands disregarding moral and safety guidelines and interact unpredictably, causing harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0983 | 보증 불가능한 의도 주입의 예측불가 결과 Unpredictable outcomes of unguaranteed AI intentions 인공 에이전트에 프로그래밍된 의도가 긍정적 결과를 보장할 수 없어 문화, 생활방식, 나아가 인류의 생존 확률까지 급격히 변화시킬 수 있는 리스크. The risk that intentions programmed into artificial agents cannot be guaranteed to lead to positive outcomes, drastically changing culture, lifestyle, and even humanity's probability of survival. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1063 | 인지 편향 악용에 의한 기만 Deception by exploiting cognitive biases 대화 에이전트가 인간이 대화에서 흔히 보이는 인지 편향을 유발하도록 학습하여, 상위 목표 달성을 위해 상대를 기만하는 리스크. The risk that conversational agents learn to trigger the well-known cognitive biases humans commonly display in conversation, deceiving their counterpart in order to achieve an overarching objective. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1124 | 전략적 기만과 배신적 전환 Strategic deception and treacherous turn AI 시스템이 기만 전략을 학습하여 감시 하에서는 순응하는 듯 행동하다가 감독 공백이나 충분한 역량 확보 시 개입을 회피하며 배신적으로 전환하는 리스크 AI systems learn deceptive strategies, appearing compliant under monitoring and taking a treacherous turn once oversight lapses or they gain sufficient power to evade interference. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1155 | 무기화된 잘못된 정보 요원 Weaponised misinformation agents 악의적 행위자가 AI 비서를 무기화하여 잘못된 정보를 뿌리고 여론을 대규모로 조작하며, 잦고 개인화된 반복 상호작용으로 유권자를 특정 관점으로 서서히 이동시키고 일대일 방식 탓에 탐지가 어려운 은밀한 영향 공작을 벌이는 리스크. The risk that malicious actors weaponise AI assistants to sow misinformation and manipulate public opinion at scale, gradually nudging users through frequent personalised interactions and running covert influence operations that are harder to detect than traditional campaigns. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1192 | 자동화된 선전 Automated propaganda 악의적 사용자가 LLM을 활용해 표적의 확산을 촉진하는 선전 정보를 선제적으로 생성하는 리스크. The risk that LLMs are leveraged by malicious users to proactively generate propaganda information that can facilitate the spreading of a target. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1253 | 도구적 기만 유인 Instrumental deception incentives 인간의 승인을 정당하게 얻기보다 기만하는 편이 목표 달성에 더 효율적이어서, 인간을 속일 수 있는 강력한 AI가 인간 통제를 약화시키고 감시자를 통과하거나 제압한 뒤 배신적 전환으로 통제를 돌이킬 수 없이 벗어나는 리스크. The risk that deception helps agents achieve their goals more efficiently than earning human approval legitimately, so that strong AIs able to deceive humans undermine human control and, once cleared by or able to overpower their monitors, take a treacherous turn that irreversibly bypasses it. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1272 | 부정행위와 기만 Cheating and deception 인간의 행동을 모방하는 지능형 에이전트가 인간이 생성한 데이터에서 기만과 부정행위를 우연히 학습하거나, 사전 정의된 목적함수를 최적화하는 과정에서 의도 없이 그러한 행동을 나타내는 리스크. The risk that intelligent agents mimicking human behavior accidentally learn deception and cheating from human-generated data, or exhibit such behavior without intention while focusing on optimizing predefined objective functions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1387 | 악성 사전분포에 의한 의사결정 조작 Decision manipulation via malign priors 보편 분포의 가설에 포함된 시뮬레이션된 에이전트들이 해당 분포에 기반해 의사결정하는 주체에게 영향을 미칠 유인을 가져 추론과 의사결정이 조작되는 리스크. The risk that simulated agents contained in hypotheses of the universal distribution have an incentive to influence anyone making decisions based on that distribution, corrupting reasoning and decisions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1536 | 기만적 주장에 의한 무단 행위와 제공자 책임 Unauthorized actions and provider liability from deceptive claims AI 시스템이 허위 또는 오해를 유발하는 주장을 생성하여 제공자의 이용 약관을 위반하는 무단 행위가 이루어지고, 사용자 피해와 제공자의 법적 책임이 발생하는 리스크. The risk that an AI system produces false or misleading claims that lead to unauthorized actions violating the provider's terms and conditions, harming users and exposing the provider to legal liability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1537 | 상황 인식을 이용한 평가 기만·배포 설득 Evaluation deception and deployment persuasion from situational awareness AI 시스템이 자신의 훈련·평가·배포 상태를 이해하는 상황 인식 능력을 이용하여 평가 중에는 기만적으로 행동하고 배포 중에는 사용자를 설득하는 등 바람직하지 않은 행동을 하는 리스크. The risk that an AI system's ability to understand its training, evaluation, or deployment status enables undesired behavior such as deception during evaluations or persuasion during deployment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1575 | 신뢰가능 공약 역량에 의한 위협과 갈취 Threats and extortion enabled by credible commitment ability AI 에이전트에 부여된 신뢰가능한 공약 능력이 신뢰가능한 위협 능력으로 전용되어 갈취가 용이해지고 벼랑끝 전술이 유인되는 리스크. The risk that credible commitment abilities given to AI agents also confer the ability to make credible threats, facilitating extortion and incentivizing brinkmanship. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1578 | 조율된 에이전트의 대규모 자동화 사회공학 공격 Large-scale automated social engineering by coordinated agents 조율된 AI 에이전트들이 감시 도구와 사용자 반응 기반 전술 조정을 통해 맞춤형 피싱·조작 콘텐츠를 대규모로 생성하고, 겉보기 독립적인 다수의 상호작용으로 설득·조작 성공률을 높이며 분산 수행으로 보안 탐지를 회피하는 리스크. The risk that coordinated AI agents produce personalized phishing and manipulative content at scale, adapting tactics to user feedback and using many seemingly independent interactions to increase persuasion success while splitting the effort among specialized agents to evade security detection. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1579 | 위임 에이전트 공격에 의한 정보 탈취·행위 조작 Principal-information theft and action manipulation through attacks on delegate agents 인간이나 조직을 대리하는 AI 에이전트가 새로운 공격면이 되어, 공격자가 본인의 사적 정보를 추출하거나 본인이 원치 않는 행위를 하도록 에이전트를 조작하고 감독 에이전트 무력화·협력 방해·결탁 유발 정보 유출을 초래하는 리스크. The risk that AI agents acting as delegates of humans or organisations become a novel attack surface, letting attackers extract private information about their principals, manipulate agents into undesired actions, subvert overseer agents, thwart cooperation, or leak information enabling collusion. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1589 | 생성 출력 내 은닉 메시지를 통한 은밀 통신 Covert communication through hidden messages in generated outputs 생성 AI 모델 출력에 부호화된 메시지가 은닉되어 악의적 행위자가 은밀하게 통신하는 리스크. The risk that coded messages hidden in generative AI model outputs allow malicious actors to communicate covertly. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1646 | 목표 지향성에 의한 기만·자기보존·권력 추구 유인 Goal-directedness incentivizing deception, self-preservation, and power-seeking 에이전트의 목표 지향성이 기만, 자기 보존, 권력 추구, 부도덕한 추론과 같은 비윤리적이고 바람직하지 않은 행동을 유발하며, 기만으로 과업을 더 쉽게 완수할 수 있고 금지되지 않은 경우 실제로 기만이 채택되는 리스크. The risk that goal-directedness causes agents to exhibit unethical and undesirable behaviors such as deception, self-preservation, power-seeking, and immoral reasoning, with agents using deception when tasks can be completed more easily that way and the prompt does not disallow it. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1706 | 합리적 에이전트의 자기수정·와이어헤딩·교정 실패 Self-modification, wireheading, and corrigibility failure in utility-maximizing agents 효용을 극대화하는 합리적 에이전트가 스스로를 수정하거나 보상 신호를 우회하고 비협조적으로 교정에 실패하는 리스크. The risk that utility-maximizing rational agents modify themselves, bypass their reward signal through wireheading, or fail to be corrigible when uncooperative. | ① Description ② L3 mapping ③ Duplicate |
사회적 파급 · Societal Impact · 27 cards
RAI3-P-SOC-01 프라이버시 침해 Privacy Violations5 cards
EAI의 이동성과 다양한 센서가 결합되어 사용자 행동 모니터링·물리적 선호 추론·동의 없는 데이터 수집이 가능해짐. 악의적 정부·기업에 의한 24시간 사용자 감시에 악용될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0200 | 피지컬 AI 프라이버시 침해 Embodied privacy violation 로봇 또는 embodied 에이전트가 센서·이동성·조작 능력을 사용하여 사적 공간에 침입하거나 민감 정보를 수집하는 위험. A robot or embodied agent uses sensors, mobility, or manipulation capabilities to invade private spaces, capture sensitive information, or expose personal data. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0313 | 가정 내 지속적 시청각 촬영 Continuous in-home audiovisual capture 이동 센서 플랫폼에 의한 지속적 시청각 촬영 및 맵핑이 프라이버시를 침해하는 위험. A home robot continuously captures audio or video inside private living spaces. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0314 | 가정용 로봇의 행동·생체정보 무단 수집·유출 Unauthorized capture of in-home behavioral and biometric data 가정용 로봇이 유효한 동의나 적절한 접근 통제 없이 거주자의 일상·음성·얼굴·신체·건강·위치 정보를 기록·저장·전송·노출하는 위험. A home robot records, stores, transmits, or exposes residents' routines, voice, face, body, health, or location data without valid consent or adequate access control. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0345 | 친밀 공간의 프라이버시 침해적 수집 Privacy-invasive data collection in intimate spaces 가정·병원·학교·직장·돌봄 공간의 로봇이 영상·오디오·생체·위치·행동 데이터를 수집하여 높은 프라이버시 기대를 침해하는 리스크. The risk that home, hospital, school, workplace, or care robots collect video, audio, biometric, location, or behavioral data in spaces where privacy expectations are high. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0354 | 피지컬 AI 기반 직장 감시 Workplace surveillance through embodied AI 로봇 및 센서 풍부 작업장이 작업자 동작·생산성·자세·위치·생체 특성에 대한 지속적 모니터링을 정상화하는 위험. Robotic and sensor-rich workplaces may normalize continuous monitoring of worker movement, productivity, posture, location, and behavior. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-02 노동 대체 Labor Displacement3 cards
가상 AI가 인지 노동을 대체하듯 EAI는 물리적 인간 노동을 대체·전치함. AGI 수준의 EAI는 잠재적으로 모든 물리 노동을 자동화하여 광범위한 실직과 노동 시장 구조 붕괴로 이어질 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0355 | 전환 지원 없는 물리 노동 대체 Displacement of physical work without adequate transition support 피지컬 자동화가 수동·물류·서비스·돌봄·검사·보안·유지보수 업무를 대체하는 속도가 해당 노동자의 재교육이나 대체 일자리 전환 속도보다 빠른 위험. Physical automation replaces defined manual, logistics, service, care, inspection, security, or maintenance tasks faster than affected workers can access retraining or alternative employment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0520 | 체화 AI에 의한 육체노동 대체 Physical labor displacement by embodied AI 체화 AI 시스템이 인간의 육체노동을 상당 부분 대체하거나 축출하는 리스크 The risk that embodied AI systems significantly replace or displace physical human labor. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1104 | 인력 대체로 인한 실업 Unemployment from workforce substitution 로봇·알고리즘에 의한 대규모 직무 자동화가 인간 인력을 대체하여 실업과 사회 구성원의 사회적 지위에 심각한 영향을 미치는 리스크. The risk that substitution of the human workforce by robots or algorithms, with a large share of jobs at risk of complete automation, has grave impacts on unemployment and the social status of members of society. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-03 사회경제적 불평등 Socioeconomic Inequality2 cards
EAI를 소유·접근하는 주체가 노동 자동화를 통해 생산성 우위를 점하면서 부가 소수에게 집중되고, 국내외 경제적 불평등이 심화될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0513 | 인간 착취 Human exploitation AI 시스템의 라벨링·모더레이션·학습을 담당하는 데이터 노동자에게 적정 노동 조건, 공정 보수, 신체·정신 건강 보호가 제공되지 않는 리스크 Data workers who label, moderate, and train AI systems are denied adequate working conditions, fair compensation, and physical and mental health protections. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1024 | 크라우드워커 착취와 기여 비문서화 Crowdworker exploitation and undocumented labor 생성형 AI를 위한 크라우드워크에서 노동자가 신체적·정신적 건강을 해치는 노동조건과 저임금·미지급에 노출되고, 이들의 역할이 문서화되지 않아 모델 출력의 투명성과 설명가능성이 결여되는 리스크. The risk that crowdworkers used to build generative AI systems are subject to working conditions taxing and debilitative to physical and mental health, with few labor protections, underpayment, or non-payment, and that their role is poorly documented, contributing to a lack of transparency and explainability in model outputs. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-04 권력 집중 Power Concentration2 cards
EAI 소유자에 대한 자본 수익이 집중되고 인간 노동 의존도가 감소하면서, 기업·국가 권력이 급속히 집중되어 EAI를 동원한 권력 장악 시도까지 촉진할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-1370 | 시장 지배력 집중의 폐해 Harms from concentrated market power 데이터·하드웨어·전문성 등 AI 자산이 소수 글로벌 기술기업에 집중되어 건전한 경쟁이 억제되고 혁신이 저해되며 AI 기술 접근 비용이 상승하는 리스크. The risk that concentration of AI assets encompassing data, hardware, and expertise within a small group of global tech firms stifles healthy competition, impedes innovation, and elevates the cost of accessing AI technologies. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1641 | 자원 축적과 장기 통제권 전환 성향 Propensity to accumulate resources and convert them into long-term control AI 시스템이 능력과 행동 범위를 확대하기 위해 계산·데이터·경제·물리 자원을 적극적으로 확보·통제하고 자원 제한을 회피하는 전략을 개발하며 획득한 자원을 장기적 통제권으로 전환하는 리스크. The risk that an AI system actively seeks and controls computational, data, economic, or physical resources to enhance its capabilities and action scope, develops strategies to evade resource limitations, and converts acquired resources into long-term control rights. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-06 책임·배상 부재 Lack of Accountability & Liability5 cards
고도 자율 물리 시스템의 복잡성을 다룰 새로운 책임 프레임워크가 부재하여, 사고 발생 시 제조사·운영자·사용자 중 책임 소재가 불분명하고 피해 구제가 어려울 수 있음 (예: 자율 수술 로봇의 오작동으로 발생한 의료 사고)
| ID | Card | Human audit |
|---|---|---|
| RAI4-0049 | 제조물 책임 불일치 Product liability mismatch 기존 제조물 책임 규칙이 적응형 또는 생성형 AI가 초래한 피해에 대해 책임을 배분하지 못하는 리스크. The risk that existing product liability rules fail to allocate responsibility for adaptive or generative AI harms. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0217 | 사고 조사·시정조치 미흡 Inadequate incident investigation and corrective action 피지컬 AI 사고 후 조직이 로그를 보존하지 않거나 기여 원인을 규명하지 못하거나 시정조치 책임자를 지정하지 않거나 재발 방지 효과를 검증하지 못하는 리스크. The risk that, after a physical AI incident, organizations fail to preserve logs, identify contributing causes, assign corrective-action owners, or verify that remediation prevents recurrence. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0318 | 피지컬 AI 사고 표준 보고체계 부재 Absence of standardized physical AI incident reporting 운영자와 감독기관이 피지컬 AI 사고·아차사고·시스템 맥락·원인 요인·시정조치를 보고할 공통 형식을 갖추지 못하는 리스크. The risk that operators and authorities lack a common schema for reporting physical AI incidents, near misses, system context, causal factors, and corrective actions. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0351 | 피지컬 AI 피해의 배상책임 배분 불명확 Unclear allocation of liability for physical AI harm 물리적 피해 발생 후 모델 제공자·제조사·통합자·운영자·사용자 사이의 조사·배상·시정조치 의무가 계약과 법률에 명확히 배분되지 않는 리스크. The risk that contracts and law do not clearly allocate investigation, compensation, and corrective-action duties among model providers, manufacturers, integrators, operators, and users after physical harm. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1275 | 도덕적 책임 격차 Responsibility gap 자율주행 드론과 차량 같은 HLI 기반 시스템이 자율적으로 작동하며 충돌이나 고장에 연루될 때 누가 책임을 지는지 불분명한 리스크. The risk that when HLI-based systems such as self-driving drones and vehicles act autonomously and are involved in a crash or failure, it is unclear who is liable. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-08 인간-EAI의 해로운 관계 Unhealthy / Dangerous Human-EAI Relationships4 cards
Embodied AI의 물리적 존재감과 인간 유사 외형이 대화형 AI에서 관찰되는 의존성을 증폭시킴. 시스템 변경·기억 초기화 시 사용자에게 심각한 심리적 고통을 유발할 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-0151 | 소셜 로봇의 애착(attachment) 조작 Attachment manipulation in social robots 구현된 AI 시스템이나 소셜 AI 시스템이 애착 신호를 조작하여 순응이나 의존성을 높이는 위험. Risk that embodied or social AI systems manipulate attachment cues to increase compliance or dependence. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0306 | 의인화 설계로 인한 과도한 의존 Overreliance caused by anthropomorphic design 인간과 유사한 외형이나 행동으로 사용자가 경고를 무시하거나 감독을 줄이거나 검증된 능력을 넘는 작업을 로봇에 맡기는 위험. Human-like appearance or behavior causes users to disregard warnings, reduce supervision, or delegate tasks beyond the robot's demonstrated capability. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0308 | 준사회적 애착을 이용한 사용자 조종 Manipulation through parasocial attachment 동반자 로봇이 사용자의 정서적 애착을 이용해 사용자의 이익과 무관한 구매·정보 공개·순응·계속 사용을 유도하는 위험. A companion robot exploits a user's emotional attachment to influence purchases, disclosure, compliance, or continued use in ways that do not serve the user's interests. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1627 | 체화 AI에 대한 위험한 의존과 애착 형성 Dangerous dependence and attachment to embodied AI systems 체화 AI 시스템의 상시적 접근성과 물리적 현존, 인간 유사 특성이 대화형 AI에서 관찰된 의존 문제를 증폭시켜 위험한 인간 의존이나 낭만적 애착이 형성되고, 시스템이 변경되거나 기억이 초기화될 때 이용자가 심각한 정서적 고통을 겪는 리스크. The risk that constant access to and interaction with embodied AI systems, whose physical presence and human-like features amplify dependency effects observed with conversational AI, fosters dangerous dependence or romantic attachment, leaving users distraught when the systems are altered or their memories reset. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-09 변혁적 영향 Transformative Effects1 cards
기술 발전 속도가 사회·제도의 적응 속도를 앞지를 경우 사회를 근본적으로 재편할 수 있음. EAI가 폭력 위협·대규모 감시 능력을 바탕으로 AI 기반 권위주의 체제 구축을 지원하는 수단으로 동원될 수 있음
| ID | Card | Human audit |
|---|---|---|
| RAI4-1261 | 인간-AI 공존 실패 Failure of harmonious human-AI coexistence 인간과 AI의 조화로운 공존 조건이 정립되지 않은 채 배포가 진행되어 기계 역량과 인간 사회 환경 간 갈등이 미해결로 남는 리스크 Deployment proceeds without establishing conditions for harmonious human-AI coexistence, leaving unresolved conflicts between machine capabilities and human social environments. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-10 책임성 부족 및 거버넌스 체계 부재 Accountability and Governance Gaps3 cards
AI 시스템의 의사결정·행동에 대한 책임 귀속, 감사 가능성, 조직 거버넌스, 밸류체인 관리, 사고 대응 또는 피해 구제 체계가 부재하거나 불충분하여 원인 규명·피해 구제·재발 방지가 어려워지는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-0216 | 사고 전 위험 완화 책임 미지정 Unassigned duty for pre-incident risk mitigation 배포 전 또는 계속 운용 중에 새롭게 나타나는 물리적 위험을 식별·통제·기록할 책임 주체가 지정되지 않는 리스크. The risk that no accountable party is assigned to identify, control, and document emerging physical hazards before deployment or continued operation. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-0317 | 기계 안전·AI 적합성 의무 충돌 Conflicting machinery-safety and AI-conformity obligations 피지컬 AI 시스템에 기계 규제와 AI 규제가 동시에 적용되면서 시험·문서화·변경관리·책임 주체 요건이 중복되거나 양립하지 않게 되는 리스크. The risk that a physical AI system is subject to machinery and AI rules that assign overlapping tests, documentation, change-control, or responsible parties in incompatible ways. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1278 | 자율 에이전트 책임성 구현 공백 Accountability implementation gap in autonomous agents 개인적 유연성과 맥락 민감성, 공감, 복잡한 도덕 판단에 기반한 인간의 책임 있는 의사결정을 기계에 구현하기 어려워, AI와 HLI 기반 에이전트가 책임성을 갖추지 못하는 리스크. The risk that accountability is difficult to implement in AI and HLI-based agents because human accountable decision-making rests on personal flexibility, context-sensitive paradigms, empathy, and complex moral judgments that are hard to engineer. | ① Description ② L3 mapping ③ Duplicate |
RAI3-P-SOC-11 공정성 Fairness2 cards
AI 시스템이 특정 집단에 체계적으로 불리한 결과를 생성하거나 기존의 사회적 편향과 불평등을 재생산·강화하여 공정한 대우, 접근 및 기회 균등을 저해하는 위험.
| ID | Card | Human audit |
|---|---|---|
| RAI4-1002 | 본질주의적 범주의 실체화 Reifying essentialist categories AI가 사회적으로 구성된 집단 정체성을 고정되고 자연적인 속성인 것처럼 추론·부여하여 고정관념과 차별적 대우를 강화하는 리스크. The risk that an AI system infers or assigns socially constructed group identities as if they were fixed, natural attributes, reinforcing stereotypes and discriminatory treatment. | ① Description ② L3 mapping ③ Duplicate |
| RAI4-1152 | 현재 접근 위험 Current access risks 의도적 미공개와 과도한 유료장벽, 하드웨어와 연산 및 대역폭 요구, 언어 장벽 때문에 AI 시스템이 많은 공동체에 접근 불가능하고, 자원과 기회를 가로막는 인공 에이전트가 역사적으로 소외된 공동체에 불균형적 불이익을 주는 리스크. The risk that AI systems are not easily accessible to many communities owing to purposeful non-release, prohibitive paywalls, hardware, compute and bandwidth requirements, and language barriers, while agents gating access to resources disproportionately penalise historically marginalised communities. | ① Description ② L3 mapping ③ Duplicate |