SIEM Bypasses: Diffracting the “CreepyDrive URLs” Sentinel Rule

How a community analytic for POLONIUM’s OneDrive C2 gets evaded via Adversarial Detection Engineering methods.

Detecting command-and-control that lives on a trusted SaaS platform is one of the harder problems in detection engineering, because the malicious traffic and the legitimate traffic terminate at the same hostname, use the same TLS, and use the same API.

CreepyDrive, is a OneDrive-based implant associated with the POLONIUM activity cluster. It reads and writes C2 tasking and exfil through the Microsoft Graph API against a normal user’s OneDrive, so from the network’s perspective it looks like Office 365.

There’s a widely shared community Microsoft Sentinel analytic that tries to catch it CreepyDrive URLs. In this article, we unpack the detection logic and identify it’s failure modes, that is, cases where the analytic rule falures to pick up CreepyDrive activity.

The goal of Adversarial Detection Engineering is to apply a formal detection logic bug taxonomy to an analytic rule, to examine whether or not there are detection logic bugs present that could be abused by threat actors. This post assesses the rule, enumerates each logic bug, and finishes with a rewrite that closes the evasions and, more importantly, stops pretending a URL string match is a behavioral detection.

Here’s the detection logic, lightly reformatted:

https://github.com/Azure/Azure-Sentinel/blob/master/Detections/CommonSecurityLog/CreepyDriveURLs.yaml

let oneDriveCalls = dynamic([
‘graph.microsoft.com/v1.0/me/drive/root:/Documents/data.txt:/content’,
‘graph.microsoft.com/v1.0/me/drive/root:/Documents/response.json:/content’
]);
let oneDriveCallsRegex = dynamic([
@‘graph.microsoft.com/v1.0/me/drive/root:/Uploaded/.:/content’,
@'graph.microsoft.com/v1.0/me/drive/root:/Downloaded/.
:/content’
]);
CommonSecurityLog
| where RequestURL has_any (oneDriveCalls)
or RequestURL matches regex tostring(oneDriveCallsRegex[0])
or RequestURL matches regex tostring(oneDriveCallsRegex[1])
| project TimeGenerated, DeviceVendor, DeviceProduct, DeviceAction,
DestinationDnsDomain, DestinationIP, RequestURL, SourceIP,
SourceHostName, RequestClientApplication

The rule runs against CommonSecurityLog, network/proxy logs from Zscaler, Fortinet, Check Point, or Palo Alto. It fires on two things:

  1. A has_any() against two exact URLs: a data.txt under /Documents/ and a response.json under /Documents/.
  2. Two regexes matching any file under /Uploaded/ or /Downloaded/, ending in :/content.

That maps to how the implant stages data, i.e fixed tasking/response files plus upload/download working folders. The intention behind the rule is correct, yet the implementation hardcodes the sample’s incidental details as if they were invariants of the malware. They aren’t , most of them are strings the operator can change in a config file.

Let’s go bug by bug.

Bug 1: The regexes are case-sensitive & the threat actor can pick the casing

KQL’s matches regex (RE2) is case-sensitive by default (see here), and neither regex carries an inline case flag. OneDrive path addressing, meanwhile, resolves case-insensitively /uploaded/ and /Uploaded/ to reach the same folder. So this evades the regex branch completely:

graph.microsoft.com/v1.0/me/drive/root:/uploaded/malware.bin:/content

The folder still works; the regex never matches.

This only affects the two regex branches. The has_any() branch is already case-insensitive, has() and has_any() are the case-insensitive operators in KQL (the case-sensitive variants are has_cs() and has_any() has no _cs() sibling but has() semantics are ordinal-case-insensitive). So changing the casing of Documents/data.txt does not slip past the literal branch. Casing is a bypass for /Uploaded/ and /Downloaded/ only.

Bug 2: URL encoding

The rule matches literal / and : characters. Percent-encoding those (%2F, %3A) means the raw logged string no longer matches:

graph.microsoft.com/v1.0/me/drive/root%3A%2FDocuments%2Fdata.txt%3A%2Fcontent

However, a threat actor who wants to bypass the rule won’t encode slashe, they’ll likely abuse Bug 1, or Bug 3. Still, decoding is free and removes the ambiguity, so the fix does it.

Bug 3: Paths, folders, filenames, and extensions are all hardcoded

This is the one that actually matters. Everything the rule keys on is threat actor-configurable:

  1. Folder names: Documents, Uploaded, Downloaded. Can be changed to Temp, Images, Public, anything to bypass the rule.
  2. File names: data, response. Can be changed to config, update, sync, to bypass the rule
  3. Extensions: .txt, .json in the has_any() list. Can be changed to .dat, .log, .bin. (Note the regex branch already allows any filename via .*, so the extension lock only applies to the two literals)

Any of these evades:

graph.microsoft.com/v1.0/me/drive/root:/Temp/sync.dat:/content

A key observation here is that the rule has encoded a sample from a Threat Intel Report, but not a technique. The technique is “read/write file content in OneDrive via Graph.” The remediation has to generalize the path, but generalizing it naively turns the rule into a firehose, which is Bug 6.

If you are interested in learning more about SIEM rule bypass mining, please sign up to our watchlist at https://www.diffractionlogic.com/#cta

Bug 4 : Path addressing is only one way to reach a file

Graph exposes multiple addressing schemes for the same content. The rule only knows root:/<path>:/content. It’s negates:

  1. Item-ID addressing: me/drive/items/{item-id}/content, direct content access by opaque ID.
  2. Children enumeration: drive/root/children, drive/items/{item-id}/children.
  3. Special folders: drive/special/documents/…
  4. users/{id} instead of me: the same operations addressed against an explicit user object.

An implant that resolves the item ID once and then reads/writes by ID never gets caught by the root:/…:/content pattern.

graph.microsoft.com/v1.0/me/drive/items/01BYE5RZ…/content

This isn’t an obfuscation trick , it’s just a different equally normal API call. When we have path coverage this narrow in an analytic rule, it can easily be a structural gap.

Bug 5: The version is hardcoded

Every pattern pins /v1.0/. The version segment is a bypass surface, but the usual remediation, and the one you’ll see auto-generated, is to swap /v1.0// for /v[0–9]+.[0–9]+// to “catch future versions like /v2.0/.”

There is no /v2.0/. Microsoft Graph has exactly two surfaces: v1.0 and beta. There’s no v2.0, and no indication one is coming. So a [0–9]+.[0–9]+ version regex:

  • Solves a problem that doesn’t exist, and
  • Still misses the real bug, which is
graph.microsoft.com/beta/me/drive/root:/Documents/data.txt:/content

beta isn’t a \d+.\d+ string, so the “flexible version” regex skips it. If you’re going to make the version segment flexible, the alternation you actually want is (v[0–9]+.[0–9]+|beta). This is a finding that can catch a few detection engineers.

Bug 6: The two things nobody writes down

Two problems sit underneath all of the above and decide whether the rule fires at all. These are data collection assumptions.

  1. TLS visibility. RequestURL in CommonSecurityLog only contains the full path if the proxy is doing TLS inspection on this traffic. Graph is HTTPS. Without decryption the proxy sees the SNI/hostname (graph.microsoft.com) and nothing after it. Microsoft explicitly recommends bypassing inspection for the Optimize-category Office 365/Graph endpoints, and plenty of orgs follow that guidance. In those environments this rule is inert no matter how good the regex is, because the strings it matches never enter the log.
  2. has() is term-based, and the host is hardcoded. The has_any() branch tokenizes on delimiters and matches terms, not raw substrings, so its behavior against real URLs (query strings, extra path segments) should be validated. And every pattern hardcodes (graph.microsoft.com), sovereign and government clouds (graph.microsoft.us, microsoftgraph.chinacloudapi.cn) are uncovered.

How to create a more robust version of the rule without the bugs?

There are two layers here, and it’s important not to conflate them. The first restores coverage, it makes the signature robust to casing, encoding, folder/filename renames, item-ID addressing, and the beta endpoint.

The second restores precision, because a fully generalized OneDrive-content pattern matches essentially all legitimate OneDrive usage and will bury you in false positives on its own.

Layer 1: a coverage-complete pattern

/ Robust CreepyDrive / Graph OneDrive content-access signature.
// Closes: case-sensitivity, URL encoding, hardcoded path/folder/file/ext,
// item-ID addressing, users/{id} addressing, the beta endpoint, sovereign clouds.
let graphContentPatterns = dynamic([
// root:/<any path>:/content - me/ or users/{id}
@‘(?i)graph.microsoft.(com|us|de)/(v[0–9]+.[0–9]+|beta)/(me|users/[^/]+)/drive/root:/[^:]+:/content’,
// items/{item-id}/content - me/ or users/{id}
@‘(?i)graph.microsoft.(com|us|de)/(v[0–9]+.[0–9]+|beta)/(me|users/[^/]+)/drive/items/[^/]+/content’,
// children enumeration - root or items/{item-id}
@‘(?i)graph.microsoft.(com|us|de)/(v[0–9]+.[0–9]+|beta)/(me|users/[^/]+)/drive/(root|items/[^/]+)/children’
]);
CommonSecurityLog
| where DeviceProduct in (“Zscaler”, “FortiGate”, “VPN-1 & FireWall-1”, “PAN-OS”) // adjust to your feeds
| extend DecodedRequestURL = url_decode(RequestURL) // closes URL-encoding evasion
| where DecodedRequestURL matches regex tostring(graphContentPatterns[0])
or DecodedRequestURL matches regex tostring(graphContentPatterns[1])
or DecodedRequestURL matches regex tostring(graphContentPatterns[2])
| project TimeGenerated, DeviceVendor, DeviceProduct, DeviceAction,
DestinationDnsDomain, DestinationIP, RequestURL, DecodedRequestURL,
SourceIP, SourceHostName, RequestClientApplication

What changed and why:

  1. (?i) on every pattern, kills the casing bypass on the folder segments.
  2. url_decode(RequestURL) before matching, removes the encoding ambiguity.
  3. [^:]+ for the path and [^/]+ for IDs, folder/file/extension-agnostic, item IDs use a permissive class because Graph IDs aren’t limited to [A-Za-z0–9-].
  4. (v[0–9]+.[0–9]+|beta) , the version segment is flexible and includes beta
  5. (me|users/[^/]+) and the items/{id} / children patterns , covers the addressing schemes the original ignored.
  6. graph.microsoft.(com|us|de), extends past commercial cloud. Add the China endpoint host separately if in scope.

Layer 2 , make it behavioral so it’s deployable

The fix isn’t a better string, it’s turning “someone touched OneDrive” into “someone touched OneDrive in a way that looks like automated C2.” Layer the coverage-complete signature above with anomaly conditions and score, rather than alerting on the raw match, this means the robust coverage fix in layer 2 become the filter for a UEBA rule.

let lookback = 14d;
let graphHits = CommonSecurityLog
| where TimeGenerated > ago(1d)
| extend DecodedRequestURL = url_decode(RequestURL)
| where DecodedRequestURL matches regex
@‘(?i)graph.microsoft.(com|us|de)/(v[0–9]+.[0–9]+|beta)/(me|users/[^/]+)/drive/(root:/[^:]+:/content|items/[^/]+/(content|children)|root/children)’;
graphHits
| summarize
RequestCount = count(),
DistinctPaths = dcount(DecodedRequestURL),
UAs = make_set(RequestClientApplication, 10),
Dests = make_set(DestinationIP, 10),
FirstSeen = min(TimeGenerated),
LastSeen = max(TimeGenerated),
// beat regularity: low stdev of inter-request gaps ~ automation
Intervals = make_list(TimeGenerated)
by SourceHostName, SourceIP, RequestClientApplication
| extend BurstPerHour = RequestCount / ((LastSeen - FirstSeen) / 1h + 1)
// Prioritise on the signals that separate an implant from a person:
// - non-browser / scripting user-agent strings (RequestClientApplication)
// - highly regular polling cadence (compute stdev of Intervals in your env)
// - a host that has never talked to Graph before this window
// - request sizes clustered around a small fixed payload
| where RequestClientApplication !has “Mozilla” // tune: drop known-good UAs
or BurstPerHour > 20 // tune to your baseline
| project SourceHostName, SourceIP, RequestClientApplication,
RequestCount, DistinctPaths, BurstPerHour, UAs, Dests, FirstSeen, LastSeen

The tuning levers that actually distinguish CreepyDrive-style C2 from a user syncing files:

  1. Client application / user-agent. Implants rarely present a full browser UA. RequestClientApplication is often your single strongest discriminator, build an allow-list of sanctioned OneDrive/Office clients and alert those that aren’t.
  2. Cadence regularity. Polling C2 produces low-variance inter-request timing and a human’s file access typically doesn’t. Compute the standard deviation of request intervals per host.
  3. First-seen. A host or user account that has never hit the Graph drive API suddenly doing so is worth more than one that does it all day.
  4. Payload size clustering. Tasking/response files tend toward a small, stable size band.

Deploy Layer 1 for coverage, gate it with Layer 2 for precision, and before either, confirm you actually have TLS-inspected visibility into Graph traffic, or none of this sees anything.

If you’re interested in learning more about SIEM Bypasses, Detection Logic Bugs, and robustness of your detection rules, be sure to check out https://www.diffractionlogic.com for the upcoming Certified Adversarial Detection Engineer training course, or https://adeframework.org for the full Detection Logic Bug taxonomy, with examples.

Takeaway

When you’re detecting C2 that lives on a trusted service, the string is never the invariant. The behavior, automated, periodic, non-browser access to a small set of files by an account that shouldn’t be scripting against Graph, that it usually invariant. Write the signature so it can’t be trivially renamed around, then let behavioral analysis decide what’s worth an analyst’s triage time.

SIEM Bypasses: Diffracting the “CreepyDrive URLs” Sentinel Rule was originally published in Detect FYI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Introduction to Malware Binary Triage (IMBT) Course

Looking to level up your skills? Get 10% off using coupon code: MWNEWS10 for any flavor.

Enroll Now and Save 10%: Coupon Code MWNEWS10

Note: Affiliate link – your enrollment helps support this platform at no extra cost to you.

Article Link: Medium