Back to Rankings

ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models

Hongyu Jin, Wenda Zhang, Runqiu Fei, Gongping Huang, Mike Conway, Ting Dang

Sep 14, 2026arXiv:2609.15089v1
cs.AI
Share
Scorecard· 16/16
4.5/10 impact

Timely, well-motivated psychology-grounded framing with a useful dataset and metrics, but undermined by synthetic references, inconsistent gains across backbones, and a shallow rule-based prompt-level mechanism.

Abstract

Empathetic response generation in spoken dialogue systems requires both accurate emotion perception and appropriate emotion regulation. Grounded in psychological theories such as the Perception-Action Model and emotion regulation theory, effective empathy depends not only on inferring a user's affective state but also on regulating how it is expressed in responses. However, recent large audio-language models (LALMs) largely treat emotion as a direct conditioning signal, lacking explicit regulatory mechanisms, which often leads to affect mirroring rather than calibrated support. We propose ER-EDF, a psychology-grounded framework that explicitly decouples emotion perception and emotion regulation in LALMs. Perception tracks the user's emotional state, while regulation determines how this state should guide empathetic response generation. The framework is model-agnostic and integrates seamlessly into existing LALMs. We further construct a spoken empathetic dialogue dataset and introduce empathy-aware evaluation metrics beyond lexical matching. Experiments across five LALMs and two datasets show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.

AI Impact Assessments

(1 model)

Scientific Impact Assessment

Core Contribution. ER-EDF introduces a psychology-grounded framework that explicitly separates *emotion perception* from *emotion regulation* in large audio-language models (LALMs) for empathetic spoken dialogue. The central claim is that existing systems treat emotion as a direct conditioning signal, causing "affect mirroring" rather than calibrated support. The proposed solution operationalizes Gross's Extended Process Model of emotion regulation by (1) fine-tuning an LALM for speech emotion recognition, (2) mapping the predicted categorical emotion into a discrete valence-arousal state, (3) selecting one of four rule-based regulation strategies (de-escalate, validate, share joy, maintain), and (4) injecting that strategy plus an emotion trajectory into a prompt for a frozen backbone. Secondary contributions are a synthetic spoken empathetic dataset (9,384 turns built on IEMOCAP/MELD) and three empathy-aware metrics (ECI, Emo-BERT, Emp-BERT).

Methodological Rigor. The design is reasonable but has notable soft spots. The core mechanism is essentially prompt-level control — the regulation strategy is a hand-crafted heuristic mapping, not a learned module (the authors acknowledge this). The dataset's "reference" responses are themselves generated by another LLM (a fine-tuned LLaMA empathy model), so the automatic metrics measure similarity to synthetic targets rather than to human gold standards. Evaluation across five LALMs and two datasets is commendably broad, and the ablation (SER-only, emotion-no-regulation, full, oracle) is well-constructed to isolate the regulation component's marginal contribution. However, results are inconsistent: ER-EDF loses to baseline on Phi-4-MM (both datasets) and on three of five backbones on IEMOCAP. The human evaluation is small (814 judgments, 77 instances, 11 reviewers) with agreement rates frequently below 50% on IEMOCAP, and no statistical significance testing is reported. The honest failure-case analysis (over-regulation, speaker-role drift, weak grounding) is a strength that partially offsets these concerns.

Potential Impact. The application area — affect-sensitive spoken dialogue for mental health, eldercare, and well-being coaching — is genuinely important and growing. The model-agnostic, inference-time nature of the framework lowers the adoption barrier. The conceptual reframing ("empathy as regulation, not just perception") is a useful contribution that could shape how the spoken-dialogue subfield thinks about empathetic generation. However, the inconsistent empirical gains, reliance on synthetic references, and simplistic rule-based policy limit the likelihood that this specific framework becomes a durable standard. Its influence is more likely to be as a motivating framing plus reusable dataset/metrics than as a widely adopted system.

Timeliness & Relevance. Very timely. LALMs (Qwen2.5-Omni, MiniCPM-o, Phi-4-MM, LLaMA-Omni) are an active frontier, and empathetic/emotional capability is an emerging bottleneck with clear commercial pull. The bridge to affective science is well-motivated and current.

Strengths.

  • Clear psychological grounding with explicit theoretical citations.
  • Broad model coverage (five recent LALMs) and a proper ablation isolating the regulation module.
  • Contribution of a dataset and empathy-specific evaluation metrics beyond lexical overlap.
  • Transparent reporting of failure modes and mixed results, and a triggering-rate analysis showing the mechanism activates selectively.
  • Limitations.

  • Synthetic gold references undermine the validity of automatic metrics.
  • Rule-based, non-learned regulation policy; the whole system is effectively prompt engineering over a frozen backbone.
  • Inconsistent results across backbones; some baselines and prior systems (BLSP-Emo) match or beat it.
  • Small human study without significance testing.
  • Limited to English acted (IEMOCAP) and sitcom (MELD) data — questionable transfer to real distressed users in deployment settings, which is the stated motivation.
  • Lexicon-based metrics (WWBP) carry known coverage biases, which the authors themselves invoke to explain outliers.
  • Additional observations. Reproducibility is moderate: prompts are fully given in the appendix and all models/datasets are public, but no code release is mentioned and the synthetic dataset construction depends on a specific external model. The valence-arousal mapping tables are provided, aiding replication. The paper is well-written and logically organized. The framing that emotion perception alone is insufficient is a mild challenge to the prevailing LALM design pattern, though not a refutation of a load-bearing empirical claim. Overall this is a competent, timely, and well-motivated contribution whose impact is tempered by evidentiary weaknesses (synthetic references, mixed gains, small human eval) and a mechanistically shallow (prompt-level, rule-based) core method.

    Rating:4.5/ 10
    Significance 4.5Rigor 4Novelty 5.5Clarity 6.5

    Generated Sep 15, 2026

    Comparison History (0)

    No comparisons yet.