Anyone who spends time online needs to understand how online content moderation is always changing. When big social media sites like Meta talk about “transparency,” they mean that they should show you everything that is being taken down and everything that is still live. Meta‘s latest Integrity Report for Q1 of 2025, which says that moderation errors in the U.S. have gone down by 50%, is certainly interesting, but this number needs to be looked at more closely before celebrating.
This change fits with Meta’s strategic choice to let its users help with some of its content moderation through the Community Notes system. The main idea is to give people the tools to spot false information, so that moderation can be based on what everyone agrees on. At first glance, this seems like a democratic and user-centered way to do things. But the trade-offs might not be the best.
The “Myth of the Perfect Note” is a big problem with the way Community Notes is set up. The system has grown to include things like Reels and Threads replies, which let users ask for notes themselves. However, there is still a major problem. The need for a wide range of users to agree on something across the political spectrum means that content that is inherently divisive is less likely to be flagged correctly. Data from similar systems clearly shows this problem: a shocking 85% of Community Notes never show up on posts. This brings up an important question: what good is a system where most contributions go unnoticed? The reported rise in content moderation metrics could actually mean that enforcement is going down. This would mean that there would be fewer attempts to actively deal with and fix hate speech, misinformation, or abuse, rather than a real drop in mistakes.
Meta’s report for the first quarter of 2025 shows some big changes in content trends. People have noticed that there is more nudity and sexual content on Facebook. Also, Instagram saw more content that was flagged as related to dangerous groups. This was reportedly because of a bug. On a more positive note, Instagram’s spam removal got better, probably because detection tools got better. Sadly, content about suicide, self-harm, and eating disorders also went up. While these trends may be partially attributed to improved detection capabilities via advanced language models, they simultaneously highlight a more disconcerting reality: harmful content persists in being generated and, in numerous cases, evades moderation efforts.
Adding to these worries are specific drops in automated detection. The report says that the automated detection of bullying and harassment has gone down by 12% and the proactive flagging of hate speech has gone down by 7%. Meta says that these drops are because they are trying to cut down on “false positives.” But this change makes us question what it really costs. When algorithms are less strict about who they let in and users have to make up for this lack of algorithmic vigilance, the whole moderation process becomes a risky bet. It could take thousands of users seeing harmful content before one person decides to flag it, and even then, there’s no guarantee that it will be acted on.
Even though these trends are happening, Meta’s internal tools are definitely getting better. Reports say that Large Language Models (LLMs) are better than human moderators at certain types of content. Still, the overall data points to a strategic change: a move away from proactive moderation and toward reactive, crowd-based governance. This strategy might not make the internet a safer place; instead, it could be a quick way to make things easier for users.
Facebook’s platform, on the other hand, is still a hard place to find good content. In the first quarter of 2025, an amazing 97.3% of Facebook posts in the U.S. did not include a link to outside content. This means that most of the time, users stay within Meta’s ecosystem, where the company has full control. This internal focus makes a space that is being filled more and more with AI-generated content that doesn’t take much effort, speculative stories, and viral, often shallow, material. There was only a small 0.6% rise in posts with outbound links, but the platform is still a tough place for good publishers to get noticed. Some anecdotal evidence might point to some improvements in traffic referrals, but the data shows that it is still a long way to go.
A 50% drop in moderation mistakes looks like progress on the surface. But there are a lot of things that affect the story, like fewer reports, smaller detection systems, and a Community Notes system that isn’t very visible to the public. Because of these things, the full story is still mostly untold, hidden in what isn’t being reported. Even small changes in enforcement have a big effect on Meta’s platforms because they are so big. A few percentage points in these changes mean millions of posts. Meta’s platforms are more than just social apps; they are global communication centers. A 7% drop in hate speech detection or a 12% drop in bullying and harassment detection adds up quickly, especially for online communities that are already at risk.
It seems like a good idea to give users the power to moderate content. But the digital ecosystem doesn’t get safer if the tools given aren’t clear, if it’s almost impossible to get a wide range of people to agree, and if automated detection systems are purposely put on the back burner. Instead, it just gets quieter, which is worrying. The most important question is not whether there are fewer mistakes. The more important question is: how many major issues are being completely ignored? Meta may be less vigilant at the very time when it needs to be more so in order to seem more open and “community-led.”