Table of Contents
The attachment scanned clean. One page, A4, 8,692 bytes, PDF 1.5. No JavaScript, no URI, no submit action, no auto-action. Its properties were emptier still: author anonymous, creator anonymous, title untitled. By the two checks an analyst usually runs on a suspicious PDF, this file had nothing left to say.
Inside those same 8,692 bytes, in a compressed stream no text-extraction pass can read, the document carried a readable signed provenance manifest that names its generator.
Two Metadata Layers In One File, Telling Different Stories
The message reached an international advertising group earlier this month, addressed to a senior advertising-operations lead. The display name claimed a national tax authority and the subject offered new taxation account information for a 2026 income statement. The body was one hand-styled HTML notice, unpersonalised: no recipient name, no account number, no reference number. That pretext is the least interesting thing here.
A properties dump of the real bytes returns that Info dictionary plus a producer string of 3.0.38 (5.1.24), a creation date of August 14, 2026 at 18:36:53 UTC, and a modification date twenty-five days later. Why the dictionary is blank is not recoverable, and intent is not the point. The verified fact is that this file has a second metadata layer, and the two layers disagree.
What the Signed Layer Records
One object is a FlateDecode stream of 5,116 bytes, over half the file. Inflated, it is a JUMBF box structure holding a C2PA Content Credentials manifest store at specification version 2.2.0, containing:
- An actions assertion whose only action is
c2pa.created, with a digital source type oftrainedAlgorithmicMediafrom the IPTC vocabulary. - A software agent recorded verbatim as the string
gpt-5-5. - A claim generator name of
ChatGPTand a document title ofimage.pdf. - A hard binding of type
c2pa.hash.data, a hash over the file with the manifest box excluded. - A signature box carrying an X.509 chain whose leaf subject reads
OpenAI OpCo, LLC, issued bySSL.com C2PA ICA R1.
The manifest's recorded creation time is August 14, 2026 at 18:36:54 UTC; the sandbox's independently run properties dump, on bytes with a matching MD5, reports a creation date one second earlier. Two readers of the same file, two unrelated encodings, one moment.
Three details rule out the objection that this manifest rode in on an embedded image. The only action is c2pa.created, not opened, placed or edited. There is no ingredient assertion. And the hard binding covers the whole file rather than a component of it. The manifest describes this document.
See Your Risk: Calculate how many threats your SEG is missing
A Free-Text Producer String Is a Hint, a Signed Manifest Is Evidence
Say the limitation plainly. This is a readable signed provenance manifest that names its generator. It is not a validated signature. Nobody verified the chain cryptographically, and the hard binding almost certainly no longer holds: the file was re-saved at 00:06:44 UTC on the morning of the send, two hours and twenty-two minutes before it left the host and twenty-five days after signing, and nothing re-signed it. A validator run against this file today would most likely report the manifest as tampered, not valid.
That is the point rather than a caveat. The naming survived a re-save that broke the math around it, because whoever handled the file never stripped the manifest. Producer and creator strings are free text that any tool writes and any operator can rewrite, which makes them a hint and nothing more. A manifest is bound to the bytes and signed under a certificate chain, so it identifies the pipeline that emitted the document even when the binding is stale. Here the trivially editable layer was blank and the signed layer was the only field that was neither empty nor false.
Who re-saved the file is not recoverable, and this post does not guess. But a document made once and mailed weeks later behaves like inventory, not like a lure improvised for one target.
Everything the Message Said About Itself Was Checkable and Wrong
The body insists twice that the attachment is password protected and that the reader should use "2026" to open it. The file is not encrypted: one properties call reports encryption as false, and the sandbox's attachment analysis agrees. The rendered page then tells the reader to use the sharing link below, consumer cloud-storage language inside a tax-authority document, pointing at a link the document does not contain. The filename performs the same mismatch.
The only clickable path in the message is one body anchor labelled Download PDF, on a subdomain of an unrelated .com whose ownership no lookup resolves. It redirects to a nonsense-concatenation .info domain that stops at a bot-verification interstitial. The landing content was never observed, and that is where the evidence stops. Credential theft is the model's label, not an observation of a page.
Authentication Passed Because Nobody Ever Published a Policy
The operator submitted the message through a European web-hosting provider's authenticated SMTP as a customer mailbox on an unrelated small business's own domain, so that host, the sender itself rather than a relay, signed it on the way out. The domain publishes no SPF record at all, so the receiving provider evaluated spf=none rather than a failure, and DMARC passed on the DKIM leg alone under a policy of p=NONE. Under DMARCbis that is a conforming pass, achieved as the true sender, while the display name claimed an agency whose domain appears nowhere in the headers. The address in the To line was on an Australian domain in the same corporate group as the protected mailbox, which is all the headers show.
Themis scored the message at 88 with labels for credential theft and a VIP recipient, against a first-time sender with a high risk level and no prior correspondence in either direction. The brand-impersonation flag never fired: a national revenue agency in a display name is not in a VIP or brand directory the way an executive or a major software vendor is. The catch came from behaviour and content, not from a list. The incident auto-resolved as phishing, with no mitigation action recorded.
Two queries follow. Alert on inbound mail that passes DMARC with spf=none, meaning DKIM-only alignment from a shared hosting gateway, a far cheaper hunt than chasing authentication failures. And add a provenance read to attachment triage: dump the Info dictionary, then check for a manifest. The 2026 Verizon Data Breach Investigations Report puts the human element in 62% of breaches and phishing as the initial access vector in 16%. Both CISA's phishing guidance and the NIST definition of phishing put verifying a sender's claims ahead of judging how a document looks. Attachment provenance is now two layers deep, and the layer that is trivial to blank is the one everybody reads.
Indicators of Compromise
| Type | Indicator | Context |
|---|---|---|
| Domain | adjusteddocuumentedrationsda[.]info | Attacker-controlled landing domain; nonsense concatenation, not a lookalike |
| URL | hxxps://adjusteddocuumentedrationsda[.]info/ | 302 destination; HTTP 403 behind a bot check, never rendered |
| URL pattern | An xd. subdomain on an unrelated .com, anchor text Download PDF | Single body anchor and redirector; ownership unresolved |
| File | september_statement_onedrive_shared_file111489 (1).pdf | 8,692-byte single-page A4 PDF; platform verdict clean |
| Hash (MD5) | 5d5dda4602f50a1bffa95502ad187b1b | Attachment |
| Hash (SHA1) | 084c32583b31c6e086b4dd49a649eeb561aa9978 | Attachment |
| Hash (SHA256) | dca49302be479156654a252ca457c608ad3bdeb31d6839384414ff1a02ce0cac | Attachment |
| Metadata | Info dictionary reading anonymous, anonymous, untitled, producer 3.0.38 (5.1.24) | Blank forgeable layer |
| Metadata | Manifest strings c2pa.created, trainedAlgorithmicMedia, gpt-5-5, image.pdf | Signed layer in a compressed stream; invisible to text extraction |
| Sender pattern | Authenticated submission as a customer mailbox at a shared web host, no SPF record, p=NONE | Hunt as DMARC pass with spf=none |
| Body claim | Attachment described as password protected, password 2026 | The file is not encrypted |
MITRE ATT&CK Mapping
- T1566.001, Phishing: Spearphishing Attachment: a one-page PDF sent as a tax statement, clean by static inspection.
- T1566.002, Phishing: Spearphishing Link: one body anchor through a redirector to a gated landing domain.
- T1684.001, Impersonation: Legitimate Organization: a revenue agency reproduced in display name, branding and copyright line, unrelated to the sending domain.
Related attacks
| Attack | What happened |
|---|---|
| The B2B Content Marketing Email That Borrowed a Brand, a Relay Allow-List, and a Security Vendor's Own URL Wrapper | A polished B2B research report offer used SelectHub branding, passed through an allow-listed mail relay at SCL -1. |
| The Fake IRS Notice That Disproved Itself | A forged IRS notice reached a commercial construction contractor with a formal detail table naming a tax year whose return has not been filed. |
| The Punycode Host That Imitated Nothing | A patient-portal reward lure pointed at an internationalized-domain subdomain. |
| Microsoft Bookings as a Weapon: When DMARC Says Trust Me and ARC Quietly Disagrees | A phishing email sent from bookings.microsoft.com passed every authentication check. |
| SAM.gov CAGE Code Scam Passes Every Auth Check | A fee-solicitation scam impersonated a mandatory SAM.gov registration requirement, put the target's real CAGE code in the subject line, and passed SPF. |
Explore More Articles
Say goodbye to Phishing, BEC, and QR code attacks. Our Adaptive AI automatically learns and evolves to keep your employees safe from email attacks.