Risk & safety · Research field
AI alignment
Catalogued:
What it means
Research and engineering concerned with making system behavior match intended goals and constraints. The term covers different technical problems and different ideas about whose intentions matter. A model following a user’s request is not automatically behaving in a socially acceptable way.
The strongest case
More capable systems can cause larger failures when objectives are misspecified or oversight is weak. Studying unwanted behavior, generalization and control before deployment can prevent avoidable harm.
The difficult part
Whose values count, how conflicts are resolved, and whether observed compliance persists in new situations remain hard questions. Alignment also cannot substitute for access control, accountability or ordinary product testing.
Sources & further reading
Sources support the factual context. The jokes are our responsibility.