Skip to content

Those who look at the worst so that AI learns to refuse it

Part two. For a model to refuse the unspeakable, humans had to show it to them. Who they are, what it costs them, and where things really stand in 2026.

Advertisement
Warning and method 🛡️
This article deals with occupational exposure to violent content. We do not describe any such content, and we stick to the categories as they appear in public court documents. If you work in this kind of job, or have done so in the past, some passages may be difficult to read. This piece is the second part of our series, following our article on the invisible workers of AI.

In the first part, we described a broad population: those who label data to train models. An attentive reader pointed out that we were conflating two very different realities. They were right, and that distinction deserves an article of its own.

Two jobs with nothing in common

Labelling images of road signs, transcribing recordings, comparing two model responses to say which is better: that is the bulk of data work by volume. It is repetitive, often poorly paid, but carries no particular psychological risk.

Then there is the other work, far smaller in numbers and completely different in nature. The work of identifying, categorising and documenting the most violent content that exists, so that a system can learn to recognise it.

The two populations sometimes overlap, because the same subcontractors move from contract to contract and the same people shift from one assignment to another. But conflating the two, as we somewhat hastily did, amounts to diluting the latter into the former. Yet it is the latter that poses the more serious problem.

Why this work exists, and why it is hard to avoid

Here is the mechanics, and it is relentless. We explained in our article on training that a model learns its behaviour from human judgements. For an AI to refuse to produce abhorrent content, it must be able to recognise it. And for it to recognise it, someone must have shown it to it, labelled it, categorised it.

In other words: the politeness with which an assistant politely declines an odious request is not an emergent property of the model. It is the product of human work done upstream by a person who looked, so that millions of others would not have to.

It is the same principle as the guardrails we described elsewhere: they do not fall from the sky, they are shaped by hand.

What the court documents say ⚖️
The lawsuits filed in recent years, first against platforms and then against AI players, describe repeated exposure to categories that we will merely name: child sexual abuse, extreme violence, acts of self-harm and suicide. The plaintiffs allege post-traumatic stress disorder, anxiety and depressive disorders. Two more recent notions appear in proceedings targeting AI: moral injury, which describes the effects of having acted against one's own values, and institutional betrayal, which describes the feeling of abandonment when the employer fails to protect.

Is this still the case in 2026?

That is the question you are probably asking, and the honest answer is: yes, with nuances.

What has changed. The issue has come out of the shadows. A class action against a major platform ended in a settlement of several tens of millions of dollars for subcontractor moderators. Protections now exist at several direct employers: psychological support, blurring by default, task rotation. And above all, proceedings now specifically target AI training work, not just social media moderation. Before a California court, a group of contractors is suing a large annotation company over the psychological harm linked to processing violent content intended to train models.

What has not changed. The subcontracting architecture remains, and it still produces the same effect: it lets clients keep operational control while maintaining legal distance from working conditions. An international union report notes a pattern that raises questions: after poor conditions were exposed in one country, a client moved its contracts to another country where pay and conditions were reportedly even worse. The problem is not being solved, it is being relocated.

Another documented finding: at equal exposure, in-house employees and contractors develop comparable disorders. What differs is the protections and recourse, not the vulnerability.

What solutions exist, and they are anything but utopian

This is the most important point of this article, because it is constructive. Protocols have been formulated by international union organisations together with the workers concerned. They boil down to a few very concrete measures.

Cap daily exposure. There is a duration beyond which the brain no longer adapts, it simply absorbs. Limiting it is the most effective measure.

Remove unrealistic performance quotas. Pressure on volume prevents any recovery break and turns every piece of content into a production line.

Psychological support available continuously, including after leaving. The protocols mention at least two years after the end of the contract, because symptoms often appear with a delay.

Train supervisors, guarantee decent pay and the right to unionise.

The argument put forward by the general secretary of an international union confederation deserves attention, because it usefully shifts the debate: exposure to distressing content may be inherent to the job, but trauma does not have to be.

The parallel she draws is hard to refute. Emergency workers, police officers, war reporters are also exposed to horror by the nature of their role. These sectors have long developed proven protocols: debriefing, mandatory follow-up, recognition of occupational risk. No one claims a firefighter will never see anything hard; what is organised is what comes after. Nothing justifies this not being the case here.

Can AI handle this work?

That is the hope often put forward, and it is partly justified. Automated systems now filter the vast majority of the most obvious content, which genuinely reduces the volume submitted to human eyes.

But two limits remain. On the one hand, it is precisely the edge cases that get escalated to humans: those the machine cannot decide, so often the most ambiguous and the most taxing. On the other hand, automation itself relies on examples labelled by humans, and every new form of content, every new workaround requires new labelling. The loop never fully closes.

Automation reduces the volume. It does not remove the function, and it concentrates the difficulty.

What you can do about it

There is no spectacular individual action to recommend here, and it would be dishonest to invent one. But there are three reasonable things.

Know. That is already a lot. An invisible job is a job without bargaining power. The mere fact that this work is named, documented and discussed has produced more change in five years than in the previous twenty.

Support recognition of occupational risk. That is the most powerful lever, because it turns a question of corporate goodwill into a legal obligation. It is exactly what Kenyan workers are demanding in court, and what the draft policy we described in the first part provides for.

Do not stop at outrage. The companies that have progressed most are not those that were denounced the most, but those that accepted verifiable commitments. Useful criticism is criticism that demands protocols, not just apologies.

One final remark. We regularly write here about what AI changes, what it costs, what it promises. This article is a reminder that part of what we value most in these tools, namely their quiet refusal to produce the worst, has a human price that no one sees. That is not a reason to do without them. It is a reason to know what we owe, and to whom.

Advertisement