Every good hunt ends on a bit of a cliffhanger: you found something….now what? Is it worth a permanent detection, better kept as a hunt you re-run each quarter, or fine to write up and move on? Most of us settle it with a gut feeling or vibe. It happens in a tuning conversation that leans toward whoever has more conviction that week, and the reasoning steps tend to evaporate the moment the rule ships. The GATES Method gives you five questions, asked in any order, to make the call the same way each time.
Where it fits: between the hunt and the detection I'm big on community and building in public is what drew me to Nebulock in the first place. In the last year, the team open-sourced two frameworks that put detection engineering and threat hunting in anyone's hands, straight from the CLI, whatever your SIEM may be. The Agentic Threat Hunting Framework (ATHF) runs the hunt through its LOCK lifecycle and, at the Keep stage, nudges you to ask whether a finding could become a detection without saying how to decide. The Agentic Detection Engineering Framework (ADEF) takes over once a rule exists, documenting and maintaining it through FORGE . There's a quiet gap in the middle where the promotion decision actually happens, and that's where GATES lives. The connective piece that turns a Keep into a Find . A "gate," if you will.
The five gates GATES is a simple checklist of five questions that together spell the name. Clear all five and a hunt earns its spot as a standing detection. Miss one, and GATES still finds it a home; a Sorting Hat for detections, where every hunt gets a house.
G: Generalizable. A repeatable behavior, or a one-off?A: Additive. Does it fill a real gap in your coverage?T: Tunable. Can you tell the attack apart from normal activity?E: Exposure-tested. Did you cover the ways it can be bypassed?S: Sustainable. Can you reliably see it, and is the upkeep worth the risk?Two tiers: trust your gut, or bring evidence You can run GATES on a napkin or as a /skill in your pipeline; it's the same five questions either way. Each gate has a base version you answer from judgment in a couple of minutes, without any tools. The advanced version asks you to prove the answer with evidence you're probably already gathering: a retrohunt, an adversary emulation, a live soak. Start with the base. If your team has more resources, the advanced tier turns each gate from a judgment call into a data-driven one.
Gate
Base ( assert)
Advanced ( demonstrate)
Generalizable
Repeatable behavior or a one-off?
Cross-fleet + companion rules that generalize the technique?
Additive / Actionable
Additive: does it fill a coverage gap?
Actionable: is there a validated or automated playbook?
Tunable
Tell attack from normal? 7- & 30-day look-back: FP rate in range?
A reusable allowlist other rules can share?
Exposure-tested
Covered the bypasses, or just the obvious path?
Run emulation: did real telemetry reveal a gap?
Sustainable / Soaked
C an we reliably see it, is the upkeep fair?
Soaked live 48h–7d: performance with concurrent detections under the volume threshold?
When a finding doesn't pass A miss is really just a matter of routing. GATES gives you a justification and won't leave a finding without a determination. This method points you to the next best place to take a hunt:
Not generalizable → time-box it: a campaign indicator with a 90-day shelf life.Not additive → merge or retire it; odds are you already have this covered.Not tunable on its own → iterate on thresholds and allowlists until the signal is usable.A bypass shows up at Exposure → log it; it becomes tomorrow's hunt, or a broader coverage decision.Not sustainable → keep it as a recurring hunt, or file the telemetry gap first.Strong findings become detections, moderate ones become quarterly hunts, weak ones get time-boxed. Every finding finds a home.
350 threat hunts through the GATES We reverse-engineered the reasoning behind 350 threat hunts (Nebulock’s own plus community hunts) to both inspire the framework and define it. After a few iterations, we crosswalked it back against the same set to pressure test the method. Roughly a fifth were ready to promote, about a third became deployable after some watchlist tuning, and around four in ten read better as quarterly hunts.
What stood out was what GATES didn’t promote to a detection. The model-assisted, baseline-heavy hunts stalled at Generalizable or Sustainable .
Knowing what to keep as a hunt is as useful as knowing what to promote. Those patterns are usually far cheaper to run four times a year than to babysit as an always-on rule.
The loop starts to compound With ATHF and ADEF, the reasoning is documented in each rule, so it compounds. The next teammate, or the next agent, picks up where you left off instead of starting from zero. A bypass you caught at exposure becomes tomorrow's hunt and a promotion carries its hunt's lineage along with it, so a little less gets relearned each time around.
Once the reasoning lives on disk, you don't have to be the one walking every finding through the gates by hand. There's a name for that step up: loop engineering. Rather than prompting an agent move by move, you build the system that prompts it for you. One agent drafts the hypothesis and scores the five gates. A second agent checks that work against your rules and your current coverage, so nobody grades their own homework. Every run writes back to the rule's journal, so the next one starts a step ahead. It's a ladder, and every step is the same five questions, you choose how much of it runs by hand:
Base : Answer the five by judgment, no tooling required.Advanced : Some prompt engineering to prove each answer: a retrohunt, an emulation, and/or a soak.Loop engineering : Hand the loop to a system you designed: scheduled discovery, a maker agent and a separate checker agent, connectors that open the watchlist and file the ticket, and memory that carries the reasoning forward. You write the loop to work the gates, while you remain the engineer who reviews what it ships.The leverage moves up the ladder when you go from writing prompts to designing loops as you climb. Taken up the rungs, it starts to feel like a detection factory: tuning owned, promotion handled as code, the whole pipeline simple to run. That's the loop we run at Nebulock: agents with DE/TH context, each pass feeding the next.
Try it this week You don't need any advanced scripts to start. Pick one detection you already trust, ask it the five base questions, and log what the Exposure step surfaces as your next hunt. That's the entire loop.
GATES is open source and ships as a /skill with worked examples in the ATHF repo. ADEF should live in the same directory as ATHF so it can fully reference GATES, understand when a hunt is ready, and move it into the detection engineering lifecycle.
We built GATES for DE/TH teams who want to move faster and find conviction in their decisions. It's yours to fork, pressure test, and improve. The gap between threat hunting and detection is smaller than we were taught.