AI Detector 360

Flux Images: Why Newer Models Beat Old Detectors

By AI Detector 360 Editorial Team · · 8 min read

Darkroom drying rack under amber light holding a row of blank clipped photo prints

The question arrives in almost exactly these words: "I ran this image through three detectors and got three different answers, so which one is right about Flux?" Probably none of them, and the disagreement is the most honest signal you got. That blunt answer needs qualification, though, because it slides too easily into "detection is useless," which is not what the evidence says.

Flux AI image detection fails most often on classifiers alone, because detectors learn the artifacts of the model generation they trained on, and newer models stop producing those artifacts. Provenance data, file history and C2PA credentials give firmer ground. Compression destroys pixel-level evidence long before the file reaches you.

Key takeaways

  • Image detectors are trained on yesterday's generators, so every new model release temporarily widens the gap between what tools expect and what they receive.
  • The classic tells of hands, teeth and garbled text have largely been fixed, which makes their absence meaningless as evidence.
  • Compression is the biggest practical obstacle: a leading detector missed 7 of 10 AI images after social-media-level compression in Bellingcat testing.
  • Provenance inspection should come before any classifier, because it is cheap, explainable and occasionally conclusive.

Why flux AI image detection breaks older tools

A trained image detector is a pattern matcher with a memory. During training it sees millions of examples labeled synthetic or real, and it learns whatever separates them in that particular sample. Some of what it learns is deep, such as how light falls or how noise distributes across a sensor. A great deal of what it learns is shallow: the specific upsampling texture of one architecture, the way one model renders skin pores, a frequency-domain quirk that shows up in that family's output.

The shallow features are the productive ones during training and the fragile ones in the wild. Ship a new architecture with a different decoder, and the quirks it produced are simply gone. The detector is not wrong about the old model. It is answering a question nobody asked anymore.

This is model drift, and image detection experiences it more violently than text detection, because image architectures change more between generations than language models do. A text detector built in 2024 still finds something familiar in 2026 prose. A texture-based image detector built on 2023 diffusion output can be close to blind on a 2026 model.

How the tells changed between 2022 and 2026

Here is the honest state of the classic checklist.

TellStatus in 2022 to 2023Status by 2026
Hands and fingersReliable giveawayMostly fixed; only sloppy output fails
Text and signageGarbled nonsenseOften legible; errors are subtler
Teeth and earsFrequently deformedUsually clean
Background crowdsMelted faces, fused bodiesMuch improved, still worth checking
Reflections and shadowsInconsistentBetter, and still the strongest visual tell
Overall textureWaxy, too smoothDeliberately varied and grain-matched

Two rows still earn their place. Reflections remain hard because they demand a consistent physical model of a scene that the generator never actually built. And background crowds are hard for the same reason at smaller scale: the model is filling space plausibly rather than depicting anything.

Everything else on that list has moved from evidence to noise. Our longer walkthrough of how to tell if an image is AI-generated keeps the visual checks that survived, and it is honest about which ones expired.

The most common mistake in 2026 is treating a clean image as proof of a camera. The tells were never the point; they were only ever the easy cases. Their disappearance changes what you can conclude from their absence, and nothing else.

Compression is the quiet killer

Bellingcat's 2023 testing is the number to remember: a leading image detector missed 7 of 10 AI images once they had been through social-media-level compression. Read that as a screening rate. If you check 100 suspect images and 30 of them really are synthetic, a tool performing at that level surfaces roughly 9 of them and quietly waves through 21.

The mechanism is straightforward. Detectors often key on high-frequency information, the fine texture where generation artifacts live. Lossy compression exists specifically to throw away high-frequency information that human eyes will not miss. The compression pipeline is, by design, an artifact eraser.

Which means the single most useful thing you can do costs nothing: get the original file. A screenshot of a post is the worst possible input. A re-download from the platform is better. The file as it left the camera or the generator is best, and it is the only version that might still carry metadata.

This also reframes what a low score means. On a heavily compressed image, "no AI detected" is close to uninformative, because the conditions that would have produced a detection were destroyed before the tool saw the file. A responsible interface should tell you that, which is why our reports separate the score from the confidence attached to it rather than blending the two into one comforting number.

Is that image AI-generated?

Upload a picture and get classifier scores, provenance (C2PA/EXIF) checks and likely-generator attribution.

Try the AI image detector

Provenance first, pixels second

Most people run this backwards. They paste the image into a detector, get a number, and only then wonder where the file came from. Invert it.

  1. Check for C2PA Content Credentials. These are signed manifests describing what made the file and what edited it. OpenAI has embedded them in image output since February 2024, Adobe Firefly and Microsoft's imaging tools carry them, and Google's Nano Banana models added them in 2026. Our primer on C2PA content credentials explains what a manifest actually contains.
  2. Read EXIF and container data. Camera make and model, lens, encoder strings, timestamps. All of it is editable, so treat it as a lead rather than proof, but an image claiming to be a 2019 phone photo while carrying a 2026 encoder string has told you something.
  3. Reverse image search. Recycled real photos are still more common than fabricated ones, and this catches them in seconds.
  4. Only now, run a classifier. With multiple engines, an explicit confidence level, and a look at which regions drove the score rather than the headline percentage.

The catch, and it is a big one: platforms routinely strip provenance metadata during upload. So the outcome of step one is usually "nothing found," and that outcome means nothing found. It is not a negative result. Absence of credentials is the internet's default state, not a red flag.

What still works on the pixels

Not everything is lost when metadata is. Four things hold up reasonably well on 2026 output:

  • Scene physics under scrutiny. Shadow directions that disagree with each other, reflections that show a slightly different world, water that does not displace around objects in it.
  • Semantic impossibility. Architecture that cannot be built, text in a language that does not quite exist, jewelry that passes through skin. Generators are fluent, not grounded.
  • Statistical region analysis. Multi-engine scoring at region level rather than image level, which is what AI Detector 360's image detector reports alongside likely-generator attribution and C2PA plus EXIF inspection.
  • Consistency across a set. One suspicious image is a coin flip. Twelve product photos from one seller that share an impossible lighting setup is a finding.

The last one is underrated. Detection at the level of a single artifact is genuinely hard. Detection at the level of a pattern of behavior is much easier, and most real-world abuse comes in volume. A person fabricating one image to win an argument is close to uncatchable; an operation producing four hundred listing photos a week leaves a signature in the aggregate that no individual file carries.

Matching the check to the stakes

Picture Dan, a part-time moderator for a regional marketplace, looking at a listing for a vintage guitar with six photos that are all slightly too beautiful. He can approve, reject, or ask. A score of 71% does not tell him which, and if he rejects a real seller he creates a support ticket and an angry review. What decides it is that all six photos share a light source that does not exist in any room, and that the same seller's earlier listings used the same impossible window. That is a pattern finding rather than a detection result, and it took him four minutes.

A short framework, because the right amount of effort depends entirely on what happens if you are wrong.

Low stakes (choosing a blog header, vetting a mood-board reference): one scan, accept the uncertainty, move on. Five credits per image on our plans means a free account's 300 monthly credits covers about 60 checks, and Starter at $9.99 for 4,000 credits covers around 800. The pricing page has the rest of the math.

Medium stakes (a paid stock purchase, a marketplace listing, a contributor submission): provenance check plus multi-engine scan plus a look at the seller's other images. Document what you found.

High stakes (news publication, evidence, an accusation against a named person): no automated result should stand alone. You need sourcing, the original file, an expert eye, and a written record of your uncertainty. Our methodology page sets out how we weigh disagreeing engines, which is the part vendors usually hide.

For marketplace and licensing work, a disclosure line beats a detection score:

Seller attests this image was captured with a camera and not generated or substantially altered by AI. Where AI tools were used, they are listed here, along with the original capture file on request.

An attestation creates accountability. A percentage creates an argument.

What nobody can verify yet

There is no published, independent benchmark that measures 2026-era image detectors across current model families the way RAID measured text detectors. Vendor accuracy claims for new generators are self-reported and usually tested on uncompressed output, which is the easy condition. Treat any specific accuracy figure for a named recent model with real suspicion, including in our own marketing.

Regulation may help at the margins. The EU AI Act's Article 50 transparency obligations became applicable on August 2, 2026, requiring machine-readable marking of AI-generated content and disclosure of deepfakes. That improves provenance inside regulated distribution chains. It does not restore metadata a social platform already stripped, and it does not bind an anonymous uploader outside the EU.

The realistic position for anyone doing this work: newer image models will keep beating detectors trained on older ones, provenance will keep being the sturdier signal, and honest tools will keep telling you their confidence instead of their conviction. Our comparison of the best AI image detectors rates them on exactly that basis, and the ones that admit uncertainty tend to be the ones worth using.

Is that image AI-generated?

Upload a picture and get classifier scores, provenance (C2PA/EXIF) checks and likely-generator attribution.

Try the AI image detector

Frequently asked questions

Do the old hand-and-finger tells still work?

Much less often. Anatomy and typography were the two obvious weaknesses of 2022 and 2023 image models, and both improved substantially by 2026. You will still catch sloppy output that way, but a clean image is no longer evidence of a camera.

Does a JPEG re-save ruin a detector's chances?

It hurts a lot. Bellingcat found a leading image detector missed 7 of 10 AI images after social-media-level compression back in 2023, and images posted today usually arrive after several rounds of re-encoding. Always try to obtain the original file.

Should I trust an image detector that gives a single percentage?

Trust it as one input. A percentage with no confidence level, no explanation of which regions drove the score, and no provenance check is a guess with a decimal point attached. Ask what the tool measured before acting on it.

Is there a way to prove an image came from a camera?

Not conclusively from pixels alone. Signed provenance such as C2PA Content Credentials is the closest thing available, and only when the capture device and the whole edit chain support it. Most images on the open web carry nothing.

Sources & further reading

Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.

Related reading