Back to Rankings

Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents

Marica Notte, Ludovica Marinucci, Vieri Giuliano Santucci

Sep 10, 2026arXiv:2609.11660v1
cs.AI
Share
Scorecard· 12/16
4.0/10 impact

A timely, well-synthesized interdisciplinary framing of alignment as a developmental process, but purely conceptual with no operationalization, limiting independent influence to a niche audience.

Abstract

In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a significant limitation: the most advanced systems rely on pre-existing datasets and human feedback strategies that are powerful but insufficient in dynamic or unknown contexts. To adapt, an agent must acquire knowledge through direct interaction with its environment. One strategy to address this challenge involves introducing higher-level mechanisms, such as intrinsic motivations, which leverage curiosity and competence, to guide exploration and learning in complex environments. While this flexibility expands autonomy, it complicates the task of ensuring agents remain aligned with human goals. Alignment, already a challenge for artificial systems in general, becomes even more complex in unstructured and dynamic contexts where predefined rules prove insufficient. To be effective and adaptable, norms must be rooted in experience through an epistemological process that starting from simple, situated principles allows for the gradual construction of more complex rules through experience, autonomous learning, and cooperation with other moral agents. Similarly to children learning social norms by exploring their environment and participating in collective practices, artificial agents must also be educated toward alignment. Following Dennett, the status of a moral agent is not innate but is attributed gradually based on the ability to responsibly manage increasing degrees of freedom. From this perspective, the regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.

AI Impact Assessments

(1 model)

Scientific Impact Assessment

Paper type: Position/conceptual paper (interdisciplinary essay bridging developmental robotics, developmental psychology, philosophy of mind, and AI governance). It presents no original experiments, datasets, or formal proofs.

1. Core Contribution

The paper proposes reframing the AI alignment problem for autonomous artificial agents (AAA) as an *epistemic and developmental* challenge rather than a top-down rule-specification problem. The central thesis: norms cannot be pre-programmed exhaustively for high-autonomy embodied agents operating in unstructured environments, so alignment must be *learned* through staged, embodied, socially interactive, intrinsically-motivated development — analogous to how children acquire social norms. The paper's distinctive move is coupling this developmental-robotics framing with (a) Dennett's notion that moral-agent status is granted gradually in proportion to responsibly-managed degrees of freedom, and (b) a concrete governance suggestion: EU AI Act regulatory sandboxes reinterpreted as "pedagogical environments" where autonomy is *earned* through demonstrated alignment. The design principles — staged curricula of increasing normative complexity, autonomy contingent on demonstrated alignment, rich contextual human feedback — form the operational core.

2. Methodological Rigor

As a conceptual paper, there is no empirical or formal methodology to evaluate. The argument is constructed by analogy and synthesis: it marshals developmental psychology literature (Tomasello, Barkley), norm theory (Kelly, Andrews et al.), philosophy (Dennett, Floridi & Sanders), and intrinsically-motivated open-ended learning (Santucci, Colas, Oudeyer). The reasoning is internally coherent and the child-development analogy is deployed carefully, with appropriate caveats (e.g., explicitly flagging that human alignment tendencies may be naturally selected and thus not transferable to artificial agents; distinguishing functional moral agency from genuine moral responsibility). However, the paper stops at the level of desiderata. The four necessary conditions (embodied, intrinsically motivated, socially interactive, developmentally staged) are asserted rather than derived or validated. There is no proposed metric for "demonstrated alignment," no mechanism for how norm-internalization would be measured or how autonomy would be incrementally released, and no proof-of-concept implementation. The gap between the high-level vision and any executable research program is large.

3. Potential Impact

The paper's impact is likely to be as a *framing contribution* rather than a technical enabler. The synthesis of developmental robotics + norm acquisition + AI governance sandboxes is genuinely useful and could seed a small research agenda, particularly within the PILLAR-Robots / intrinsically-motivated-learning community from which it emerges. The regulatory-sandbox-as-pedagogy idea is a fresh and policy-relevant hook that could attract attention from AI governance scholars. However, the audience is niche, and without concrete algorithms, benchmarks, or experiments, the paper is unlikely to change practice on its own. Its citations will most plausibly come from other conceptual/position work and grant proposals rather than from methods papers.

4. Timeliness & Relevance

Highly timely. Alignment of increasingly autonomous, embodied agents is a central concern in 2024–2025, and the paper explicitly engages with very recent framing (Silver & Sutton's "Era of Experience," 2025) and the operative EU AI Act. The tension it identifies — that intrinsic motivation and open-ended learning expand autonomy in ways that complicate alignment — is a real and under-addressed bottleneck. The governance angle connecting technical alignment to concrete regulatory instruments is especially current.

5. Strengths & Limitations

Strengths:

  • Genuine interdisciplinarity: cleanly integrates cognitive science, philosophy, developmental robotics, and AI policy.
  • A crisp, memorable reframing ("agents must be *educated* toward alignment"; "autonomy is the precondition of alignment, not its threat").
  • The novel and actionable reinterpretation of regulatory sandboxes as developmental scaffolding.
  • Intellectual honesty about limits (moral responsibility, accountability, distribution of normative authority).
  • Limitations:

  • No operationalization: no algorithms, architecture, experiments, or evaluation criteria. It is a manifesto, not a method.
  • The developmental analogy, while suggestive, is not critically stress-tested — it is unclear which aspects of child norm-learning are actually implementable in RL/robotic agents versus which depend on biological/social substrates the paper itself acknowledges may not transfer.
  • Very short and shallow in technical depth; the "concrete design implications" are broad guidelines rather than specifications.
  • Limited engagement with the substantial existing technical alignment literature (RLHF limitations, inverse reward design, value learning, norm-following RL) — the alignment discussion leans on a single popular-science reference (Christian, 2020).
  • Additional observations: Reproducibility is not applicable (no empirical component). The barrier to *engaging* with the ideas is low (no compute/data needed), but building the envisioned system would require a well-resourced robotics lab. The paper does not contest any prior empirical claim; it is additive and programmatic. Its foundationality is modest-to-moderate: the framing could be reused as a conceptual scaffold, but it does not provide a reusable technical primitive.

    Overall, this is a thoughtful, timely, well-synthesized position paper whose contribution is conceptual framing rather than demonstrated results. It is likely to be cited within its niche and to inform the authors' own funded research agenda, but its independent influence on the broader field will be constrained by the absence of any concrete, validated mechanism.

    Rating:4/ 10
    Significance 4Rigor 4Novelty 6Clarity 8

    Generated Sep 11, 2026

    Comparison History (0)

    No comparisons yet.