Skip to content
dg.dev
← all projects

2026 · updated 10 Sept 2026

LLM-Fake-News-Detector

Three models, a verdict from each, and a check on whether they learned the right thing.

The problem

A university brief with an 85% accuracy target. The models cleared it easily, and that was the problem: scores that high on a static dataset usually mean a model has found a shortcut.

How it works

  1. Article text
  2. TextPreprocessor
  3. TF-IDF + LogRegBiLSTMBERT
  4. Flask: three verdicts

A ModelOrchestrator keeps a registry of three wrappers behind one interface, loads them in a background thread, and asks each one for a label and a confidence. Adding a model means adding it to the registry, not touching the logic around it.

The hard parts

  1. 01

    Finding the shortcuts

    On articles from outside the dataset, verdicts flipped on things that had nothing to do with truth. SHAP showed why: publisher names like “Reuters” and the dash characters from datelines were driving the predictions.

  2. 02

    Cleaning out the signal

    A regex strips the dateline prefix (up to 35 characters before a dash), then URLs, HTML and punctuation, before training and before inference, so the shortcut isn't there at either end.

  3. 03

    From 1.5 seconds to 10 milliseconds

    BERT is loaded once, warmed up with a dummy input and kept in memory, in a background thread so the server starts straight away. Prediction went from about 1.5 s to 8–10 ms.

Decisions and trade-offs

Three models, not one

The baseline shows what cheap features can already do, and disagreement between the three is useful information for the reader in its own right.

Each model fails on its own

Every model reports its own status. If BERT fails to load, the other two still answer and the page says which one is missing.

Whatever hardware is there

BERT picks CUDA, Apple's MPS or the CPU at start-up, so the same code runs on a laptop and on a GPU machine.

What it doesn't do yet

  • Accuracy on a static test set says little about real articles; a fresh, dated test set would be the honest benchmark.
  • The SHAP explanations live in the notebooks, not in the web app.