A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis · TMLR
Baseline: original submission of 13 June 2026. This note summarizes the substantive additions and revisions made across the review period. No headline claim from the original submission was retracted; the revisions add evidence, a second real benchmark, and presentation improvements.
The revision keeps the original contribution (a controlled synthetic ABSA corpus plus an audit–filter–validate methodology) and strengthens it on three fronts: the label-faithfulness audit is now validated against humans and independent model families, the synthetic-to-real story is completed on a second real benchmark, and every reported control is expanded to multiple seeds and a second architecture.
Paste-ready cover note to the Action Editor
Post as an Official Comment on the Action Editor thread (OpenReview renders Markdown). Per-reviewer point-by-point responses are the three boxes below.
Paste-ready summary for the OpenReview revision field
Plain text for the “List of Changes” box. Click Copy (or select all inside and copy).
Paste-ready answer for the Human Subjects Reporting field
Enter this in the submission's “Human Subjects Reporting” field. The label study used research-team annotators labeling synthetic, non-personal text, so it is not human-subjects research; all real-data benchmarks are pre-existing public datasets.
Official responses to reviewers
Post each block as an OpenReview Official Comment in that reviewer's thread (OpenReview renders Markdown).
Reviewer nfat — official response
Reviewer dWED — official response
Reviewer h7LN — official response
New empirical results
Direct human validation of the audit. A three-annotator study labels a stratified sample of the synthetic corpus for aspect presence and sentiment (Fleiss κ 0.70 on presence). Human confirmation of a declared aspect rises monotonically with the audit score, and the human labels reproduce the audit's presence-faithful / sentiment-noisier split (human sentiment agreement ≈0.40, matching the audit's strict 0.42).
Audit independence. Two independent open-weights auditors (Llama-3.3-70B, GLM-4.6) reproduce the faithfulness audit (cross-family Cohen κ 0.56–0.65), so the signal is not an artifact of a single judge family.
Provider-agnostic pipeline. Both the generation pipeline and the prompted baseline now reproduce across four generator families spanning four providers (GPT-5.4, Gemini-2.5-Flash, GLM-4.6, Llama-3.3-70B); the multi-provider zero-shot baseline lands in a single narrow band below the trained encoders.
Second real benchmark (EduRABSA) completed. Built out from a passing mention into a full multi-seed synthetic-to-real transfer table, a real-only trained reference, and a synthetic-pretrain + real-fine-tune result. Synthetic-only training now recovers about half a real-trained model on both datasets (52% Herath, 60% EduRABSA), and synthetic pre-training followed by real fine-tuning matches or exceeds real-only training on both.
Cross-architecture and multi-seed controls. The faithfulness-filtering contrasts are replicated on DistilBERT alongside BERT and extended to eight seeds, and a covariate-matched comparison isolates the filtering gain from composition and prediction-mask confounds.
Learnable-signal controls. Label-permutation collapses detection to the trivial floor (0.182 vs 0.276), accuracy scales monotonically with training size, and restricting to faithfully labeled rows raises the ceiling — establishing genuine learnable signal rather than label priors.
Truncated-row robustness. Regenerating the 841 output-token-capped rows at a raised cap restores full-corpus length adherence, and the capped rows are shown not to bias the benchmark in aspect, polarity, or faithfulness.
Presentation and structure
Abstract and introduction. Abstract completed and tightened; the contribution statement de-duplicated against Section 1.1; the paper-structure paragraph moved to the end of the introduction.
Figures. All figures redrawn on white backgrounds with enlarged, non-overlapping labels; Figure 2 redesigned with labeled panels and in-plot annotations; Figure 5's legend overlap removed; Figure 4 model names set horizontally.
Tables and cross-references. Table 1 column widths and padding fixed and inter-table spacing standardized; every table and figure is now introduced in the immediately preceding paragraph with a note on what it shows.
References. Reference list alphabetized by author surname (TMLR requirement) with all in-text citations renumbered; seven recent (2024–2026) citations added for literature completeness.
Appendices. Reorganized from a long flat list into seven thematic sections (A.1–A.7) with subsections, and a duplicated appendix table removed, so the supporting material is easier to navigate.
Artifacts. Code, datasets, the human-ranking study, and model checkpoints released as a versioned archive.
Main tables 12 → 15
Appendix tables 15 → 28
Main figures 5 → 6
Appendix sections flat → 7 themes
Point-by-point replies to each reviewer are in the author responses; the submission rationale is in the cover letter. This summary is a reading aid and is not part of the manuscript.