A classifier can read every post.
A person cannot.
That gap is why automated moderation exists, and it is also where most of the trouble starts.
A model returns a score.
Not a verdict. A number expressing how confident it is that something matches a category it was trained on.
Very high, and the system can act by itself.
Very low, and the content simply publishes.
Everything between belongs to a person, and that middle band is the whole design.
A score is a signal.
It doesn’t turn a probability into an explanation.
Which means the thresholds aren’t moral lines.
They are operating choices, and they should differ by category.
Where a missed violation carries exceptional harm — imminent violence, material involving children — the threshold sits high, and a wider false-positive rate is the price you accept on purpose.
Where a wrongful removal costs about as much as a missed one — harassment, misinformation, borderline nudity — more of the middle belongs with a human.
So don’t set one global threshold and call the system consistent.
Consistency of method can hide inconsistency of harm.
And widening a band is not free.
Lowering the bottom edge a little can double the volume arriving at human review.
That is worth doing, sometimes.
It is never worth doing without admitting what it buys and what it costs.
The queue that results has a second decision to make: not only what to review, but what to review first.
A score can raise severity. So can reach, how fast something is spreading, how many people reported it, and what the account has done before.
That is prioritization.
It is not permission to let the score finish the case.
Then give the human reviewer the material the score cannot carry.
The thread. The account’s age. What came before this post.
Sarcasm, history, a running joke, a quotation — these are exactly the things a category score has no access to, and exactly the things that decide most contested cases.
The reviewer confirms, approves, or escalates.
And the design of that moment matters more than almost anything else in the system.
Don’t put the classifier’s suggested action where it becomes the easiest thing to click.
That produces automation bias: agreement with the score rises even when the score is wrong, because overriding takes work and confirming takes none.
Then audit a sample of closed decisions, and audit it independently.
The point of the audit is to test whether the person is applying the policy.
Not to ask whether the person agreed with the machine.
When independent reviewers stop agreeing with each other, the answer is to examine the rubric, the training, or the conditions of the shift — never to make the model’s recommendation more prominent.
Some decisions shouldn’t be made alone at all.
Banning a well-known account, judging political speech, removing high-reach news: those want a second senior reviewer who concurs.
Not because one person is untrustworthy.
Because a decision that will be argued about in public should have been argued about in private first.
Then say what happened.
A notice should name the rule, say whether a machine or a person made the call, say how far the action reaches, and say where to appeal.
Whether it was automated is not a technicality.
It is the difference between a judgment someone made and a threshold someone set months ago, and a member is entitled to know which one they received.
Global removal and a regional block are not the same action either.
An explanation that hides the rule, the actor, or the scope cannot be tested.
Then read the appeals as feedback.
A reversal rate that climbs says a category is sending wrong decisions downstream — an over-aggressive threshold, vague guidance, or a first review that was too quick.
A reversal rate near zero can mean something worse: an appeal route that only confirms itself.
The point isn’t to pick a pleasing rate.
It is to let contested cases test the first decision, instead of disappearing into a closed queue.
There is one place automation genuinely finishes the job.
For known illegal material, a fingerprint match against an existing registry isn’t a probability at all. It is a lookup, and it should act at machine speed and report where the law requires.
That exception proves the rule for everything else.
Certainty is what lets a machine decide.
Everywhere else the machine is sorting, and someone still has to judge.
So use the classifier to sort work by likely risk.
Use category-specific thresholds to decide what belongs with a person.
Set the escalation route before the difficult case arrives, not during it.
State the reason, the actor, and the scope.
And read reversals as information rather than inconvenience.
The queue may get shorter.
Automation can make the queue smaller. It cannot make the responsibility disappear.
