<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Practice Guides · Bright Learning</title><description>Practice Guides. Evidence-linked steps, decisions, and limitations for educators and families.</description><link>https://brightlearning.ai/</link><item><title>When AI Cheers Instead of Critiques: A Practice Guide for LLM Writing Feedback</title><link>https://brightlearning.ai/articles/when-ai-cheers-instead-of-critiques-a-practice-guide-for-llm-ndidate4/</link><guid isPermaLink="true">https://brightlearning.ai/articles/when-ai-cheers-instead-of-critiques-a-practice-guide-for-llm-ndidate4/</guid><description>LLM-based feedback tools can change writing feedback based on stated student identity, sometimes offering warmth in place of substantive critique. Here are concrete classroom checks teachers and instructional leaders can use, paired with the honest limits of the evidence behind them.</description><pubDate>Wed, 15 Jul 2026 19:29:04 GMT</pubDate><content:encoded>&lt;h2 id=&quot;why-this-matters&quot;&gt;Why this matters&lt;/h2&gt;
&lt;p&gt;Teachers are increasingly routing student writing through LLM-based tools such as Brisk, MagicSchool, and Khanmigo to generate first-pass feedback. Many of these tools invite &amp;quot;personalization&amp;quot; based on student attributes. A recent study suggests this feature carries a risk worth naming: feedback can shift in tone and substance depending on the student identity descriptors supplied to the model.&lt;/p&gt;
&lt;p&gt;The study, titled &amp;quot;Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback,&amp;quot; examined how automated writing feedback changed when models were told about student attributes [source:src_edf81ffc5392f321]. Reporting on the work summarized the core concern in a blunt headline: AI gives more praise and less criticism to Black students [source:src_abf0264c123e695f].&lt;/p&gt;
&lt;p&gt;The practical worry is not abstract. If a tool hands some students encouragement while giving others pointed revision guidance, the students who most need substantive critique to grow as writers may receive the least of it. This guide translates the study into usable classroom checks. Before we get to the moves, a word on what the evidence does and does not establish.&lt;/p&gt;
&lt;h2 id=&quot;what-we-can-and-cannot-verify&quot;&gt;What we can and cannot verify&lt;/h2&gt;
&lt;p&gt;We read the arXiv preprint and its full HTML text [source:src_edf81ffc5392f321][source:src_eeda7aa3d489f8ad]. The final published version in the ACM Digital Library, carrying DOI 10.1145/3785022.3785113, returned an HTTP 403 error and could not be inspected directly [source:src_a667077a82e927ef]. Its existence and its April 26 publication at the LAK26 conference are corroborated only by the matching DOI and booktitle in the arXiv preprint and by the Hechinger reporting [source:src_032d05a85be5c9c7][source:src_abf0264c123e695f]. We were not able to confirm that the published text matches the preprint word for word.&lt;/p&gt;
&lt;p&gt;A key scope limit: the study tested explicit identity descriptors, not subtler cues like student names or writing samples. The paper&amp;#39;s own robustness section mentions name-only prompting, but we did not read that section in full and cannot report its findings [source:src_eeda7aa3d489f8ad]. So the practical reach of the effect to more realistic, everyday prompting is not something this guide can settle.&lt;/p&gt;
&lt;p&gt;Finally, the practice recommendations below are the editorial framing of Bright Learning and the Hechinger reporting, not formal findings of the study itself. One interview quote attributed to Tanya Baker appears in the Hechinger report and was not independently verified by us [source:src_abf0264c123e695f]. Bright Learning did not conduct original reporting or testing here; we are synthesizing published sources into classroom practice.&lt;/p&gt;
&lt;h2 id=&quot;four-classroom-checks&quot;&gt;Four classroom checks&lt;/h2&gt;
&lt;h3 id=&quot;1-turn-off-or-ignore-per-student-identity-personalization&quot;&gt;1. Turn off or ignore per-student identity personalization&lt;/h3&gt;
&lt;p&gt;If a tool lets you attach student identity attributes to a feedback request, the safest default is to leave those fields blank or turn the feature off. The study&amp;#39;s concern centers on feedback that shifts when explicit identity descriptors are supplied [source:src_edf81ffc5392f321]. Removing the input removes the most direct pathway for that shift.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope limit to keep in mind:&lt;/strong&gt; turning off explicit descriptors does not guarantee neutral feedback. The study focused on explicit attributes; it did not establish that a tool fed only a student name or writing sample behaves neutrally [source:src_eeda7aa3d489f8ad]. Treat this move as reducing a known risk, not eliminating all risk.&lt;/p&gt;
&lt;h3 id=&quot;2-keep-ai-feedback-generic-reserve-identity-aware-encouragement-for-human-judgment&quot;&gt;2. Keep AI feedback generic; reserve identity-aware encouragement for human judgment&lt;/h3&gt;
&lt;p&gt;Encouragement that responds to who a student is, their background, their history in your class, their confidence on a given day, is real pedagogical work. But it is human work. Configure the tool to critique the writing on the page and leave the relational, identity-aware encouragement to you. The reporting frames the danger as substituting warmth for critique [source:src_abf0264c123e695f]; the mitigation is to keep those two functions in separate hands.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope limit:&lt;/strong&gt; this is a design principle from the editorial framing, not a tested intervention in the study. The paper did not evaluate a &amp;quot;generic AI plus human warmth&amp;quot; workflow.&lt;/p&gt;
&lt;h3 id=&quot;3-spot-check-the-tool-strong-paper-versus-struggling-paper&quot;&gt;3. Spot-check the tool: strong paper versus struggling paper&lt;/h3&gt;
&lt;p&gt;Once a term, run a quick audit. Take one strong paper and one struggling paper, strip any identity information, and feed each through your tool. Read the two sets of comments side by side. Ask: does the tool escalate its critique for the weaker paper, as a human reader would, or does it soften? Does the stronger paper get pushed harder, or coddled?&lt;/p&gt;
&lt;p&gt;This is a coarse check, not a measurement. It will not detect subtle bias, and one comparison is not a study. But it surfaces gross tone mismatches and gives instructional leaders a repeatable routine to demonstrate the tool&amp;#39;s behavior to a team.&lt;/p&gt;
&lt;h3 id=&quot;4-read-ai-comments-for-push-versus-cheer&quot;&gt;4. Read AI comments for push versus cheer&lt;/h3&gt;
&lt;p&gt;When reviewing AI-generated feedback before it reaches a student, sort each comment into one of two buckets: does it push the argument forward (name a weak claim, demand evidence, flag a logical gap), or does it just cheer (praise effort, affirm the topic, encourage generally)? A stream of cheer with little push is a signal, especially if it correlates with which students receive it. The study&amp;#39;s central worry is exactly this pattern of praise without criticism [source:src_abf0264c123e695f].&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failure mode to watch:&lt;/strong&gt; cheer is not always wrong. Early-draft feedback and feedback to genuinely anxious writers may lean supportive by design. The problem is not warmth itself but warmth that displaces the substantive critique a student needs to improve.&lt;/p&gt;
&lt;h2 id=&quot;privacy-age-suitability-and-safety&quot;&gt;Privacy, age suitability, and safety&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Privacy.&lt;/strong&gt; Attaching student identity descriptors to a third-party LLM tool means transmitting sensitive information about minors to a vendor. Beyond the bias concern, minimize what you send. Check your district&amp;#39;s data agreement before entering any student attribute.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Age suitability.&lt;/strong&gt; These tools and this guidance apply across grade bands, but younger students are less able to recognize when feedback is hollow. The push-versus-cheer check matters more, not less, for younger writers who cannot self-diagnose thin feedback.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Safety and failure modes.&lt;/strong&gt; The main failure mode is silent: a tool that quietly delivers less rigorous feedback to some students while appearing helpful to everyone. Because the effect is invisible in any single comment, it is only detectable through comparison and routine review. Do not assume neutrality because a tool is popular or because a single output looks fine.&lt;/p&gt;
&lt;h2 id=&quot;the-bottom-line&quot;&gt;The bottom line&lt;/h2&gt;
&lt;p&gt;The evidence here is a focused lab study, corroborated at the edges and read primarily in preprint form, testing explicit identity descriptors rather than the subtler signals of real classrooms [source:src_edf81ffc5392f321][source:src_eeda7aa3d489f8ad]. That is a genuine limit. But the direction of concern is clear enough to act on cheaply: strip identity inputs, keep AI critique generic, spot-check across ability levels, and read for push over cheer. None of these moves costs much, and each one keeps the substantive critique that improves writing in front of every student.&lt;/p&gt;
</content:encoded></item></channel></rss>