AI isn’t enough to protect social media communities from AI
AI’s biases
Typical social media AI-based moderating systems use machine learning classifiers to analyze posts and identify and flag content that breaks platform rules. But it’s difficult for a machine to understand the nuances of sarcasm, satire, and slang.
Further, some research (examples here, here, and here) suggests that marginalized groups can be disproportionately affected by AI moderation. Without human oversight, AI can end up penalizing the very communities most vulnerable to the hateful content the systems are designed to combat.
Gilbert, who is also the research director of Cornell’s Citizens and Technology Lab, says that “marginalized and vulnerable populations are among those who experience the highest rates of moderation, and that typically this is a result of ‘false-positives,’” often driven by instances of counter-speech, language reclamation, and “responses to hateful content.”
“False positives are an equity issue. They mean that groups that are already marginalized are further silenced and censored,” she added.
AI moderators can also make communities less effective at moderating themselves. On Reddit, for example, some subreddit moderators would prefer to ban users who use hateful or violent rhetoric. But if Reddit’s AI removes such content before a human moderator sees it, those moderators lose the ability to assess whether a ban is warranted.
In terms of giving human mods more control, Reddit this week announced expanding testing for Rules Hub, a suite of tools that lets human mods “choose which rules should be automatically enforced, decide what happens when a rule is triggered (send to queue, filter, or remove), preview the experience before enabling it, and review logs and insights.” Reddit expects Rules Hub to eventually replace the Automod tool, which relies primarily on exact keywords.
AI is a tool, not the solution
Mods I’ve spoken with have repeatedly blamed the generative AI boom for a spike in content that breaks community-specific or broader platform rules. That’s a serious problem for social media sites that rely on user contributions.
Companies will continue to try new methods of moderating more reliably and effectively, but reducing human input is a step backward. Low-effort AI-generated content is changing the challenges moderation teams face, but that makes stronger approaches more necessary, where machine-scale detection can be combined with human judgment and expertise.
Just as social media has no value without people, content moderation can’t succeed without human judgment at the forefront.
Advance Publications, which owns Ars Technica parent Condé Nast, is the largest shareholder in Reddit.