The Framework

Chapter 1

Open Detection Engineering Framework

A detection engineering story

Mar 27, 2023

Subsections of The Framework

Introduction

Purpose

ODEF enables organizations to establish and apply fundamental principles and best practices in detection engineering. Adopting ODEF accelerates service enhancement, refines processes, and advances cybersecurity maturity.

The framework strategically aligns cybersecurity efforts with business objectives (e.g., risk reduction and operational resilience). It enables the creation of robust detection systems, provides assurance, reduces reliance on vendors, and, most importantly, strengthens an organization’s overall security stance. ODEF defines core principles for effective detection engineering and introduces a three-tiered maturity model to evaluate organizational performance.

Each phase within the framework’s core is meticulously crafted to map out the detection lifecycle, guiding detection engineers with focused functions and outcomes expectations that steer them through the entire process. The maturity levels act as strategic benchmarks, providing organizations with a high-level perspective on their detection engineering strategy and highlighting areas for potential improvement.

High-Level Goals

The overarching ambitions of the framework are to:

  • Establish a methodical, replicable, and predictable approach for developing hunts and detections.
  • Achieve extensive organizational visibility.
  • Transform insights into sustainable, actionable knowledge while fostering a culture of information sharing.
  • Promote persistent vigilance in detection practices.
  • Validate detections rigorously through systematic testing.
  • Cultivate an environment where knowledge and knowledge sharing is the main driver of security initiatives.

Framework Mindmap

“ODEF Mindmap” “ODEF Mindmap”

Framework Core

Framework Core

The Framework Core is a suite of activities aimed at achieving specific cybersecurity objectives, supplemented by illustrative examples to guide their implementation. The Core is structured around three lifecycle phases: Sunrise, Midday, and Sunset which holistically cover the detection lifecycle from inception to completion, with dedicated functions, guidelines, and goals for each phase. Together, these components provide detection engineers with a “north star” focus, enabling them to deliver high-quality detections with precision.

Functions, Goals, Guidelines

Each phase of the detection lifecycle is defined by specific functions, goals, and guidelines. These elements work together to guide detection engineers in producing robust, reliable detections

Functions

Inspired by the single-responsibility principle from software engineering, each function in ODEF represents a discrete step with a singular focus and outcome within the detection lifecycle. These single-goal activities are designed to drive a specific activity, ensuring clarity and precision at every stage.

Goals

Each function is crafted to achieve a specific, well-defined result that aligns with the overarching objectives of the detection lifecycle. While goals may sometimes be high-level or abstract, it is strongly recommended to maintain a direct correlation between functions and goals. This alignment ensures that each function contributes meaningfully toward the end goal, enhancing both focus and quality in detection engineering.

Guideline

Guidelines act as reference materials that support the achievement of goals within each function. They provide detection engineers with context-specific insights and best practices, helping to standardize processes while allowing adaptability. For example, a document detailing a company’s unique Change Management process would serve as a valuable guideline

Phases in detail

The Core encompasses three primary lifecycle phases:

These phases collectively chronicle the lifespan of a detection mechanism, from its conception to its eventual retirement/decommissioning.

Phases

Intro

In the rhythm of each day lies the strength of progress; every sunrise brings new potential, midday illuminates our purpose, and sunset reminds us to refine and renew.

Phases

ODEF pioneers the concept of a detection lifecycle, comprehensively encapsulating the journey of a detection mechanism from creation to retirement. It's structured into three distinct phases: Sunrise, Midday and Sunset. This tri-phase approach ensures thorough coverage for each stage of a detection's active life. Corresponding to each phase are specific functions, goals, and guidelines that lay the groundwork for effective and efficient detection engineering.
Mar 27, 2023

Subsections of Phases

Sunrise

Phase 1️⃣ Sunrise 🌅

Sunrise is the first phase of the detection lifecycle. It marks the inception, development and deployment of the detection. During that phase there are 6 core functions that should be addressed:

  • Research
  • Prepare (Logging)
  • Build (Detection Content)
  • Validate
  • Automate
  • Share (Knowledge)

High level goals for the Sunrise phase

  • Build high fidelity detection
  • Ensure detection validation
  • Create documentation
  • Integrate and automate in the environment
  • Socialize the detection with the security organization

FunctionsGoalDescriptionGuidelines
ResearchOpportunity IdentificationIt can be triggered from analyzing threat intelligence reports, or OSINT, or internal knowledge for a particular security gap. Document the use case and the goals of the detection as part of the opportunity identification process.
  • Document the use case that you’re building and set goals.
  • Is the TTP already covered by an existing alert or detection?
  • Is there sufficient knowledge to start building or additional research would be required?
  • What are sources of information that will assist the research?
PrioritizeDetection engineering work has to be prioritized and tracked. Work prioritization can be based on urgency and priority. Backlog of detections and security posture activities is desirable and recommended.Prioritization criteria:
  • Criticality of the system

  • Highest level of threat to the organization

  • Ease of Exploitation

  • Past incidents

Develop Research QuestionsWrite your research questions that while answering you will gain understanding of the topic.Examples:
  • Write down what you already know or don’t know about the topic.
  • Use that information to develop questions. Use probing questions. (why? what if?).
  • Avoid “yes” and “no” questions
Information GatheringResearch and collect sufficient information in order to start understanding the detectionProvides a good overview of the topic if you are unfamiliar with it.
  • Identify important facts, dates, events, history, organizations, etc. (in case the detection is a response to a past incident.)
  • Find bibliographies which provide additional sources of information (include in the Appendix section detection document)
Technical ContextCreate and understand technical context around the detection
  • Start putting technical writeup by summarizing the most important information from technical aspect
  • Research the technology associated with the technique to help understand the use cases, related data sources, and detection opportunities
  • Note: Defenders often create superficial detections because they lack an understanding of the technology involved. In case of uncertainties it is best to engage the team or engineer responsible for the management of the technology
PrepareIdentify DatasetIdentify the log source that will be used for the detectionKnow your environment
  • Understand the data source and document it by creating a data dictionary.
  • The data dictionary should grow and contain sources of data and their corresponding schemas. It can later be used to quickly refer to.
Visibility CheckEnsure there is sufficient logging, retention and visibility in order to successfully build the detection and satisfy the use case
  • Use the accumulated technical knowledge to identify source and identify the events required to build detection
  • Use any historical events in order to validate that there is sufficient visibility
Improve (optional)Once the data is explored we can identify opportunities for improvements such as:
  • Collecting additional logs or change logging levels
  • Create additional attributes (parsing of raw logs)
  • Consolidation of distinct logs
Improvement initiatives and requests should be communicated to the responsible for the dataset in question team. For that purpose it makes sense to maintain a contact list that provides quick reference to technology, support/engineering teams and contact details.
Build & EnrichDetection CreationCreate a detection query against the identified datasetHaving a good understanding of the technical context and the data source begin building queries to narrow down the data to actionable insight.
Manual TestingPerform a manual testing and ensure the query works syntax and logical perspective
  • Ensure the query does not have any syntax errors
  • In case the detection is built in response to past incident ensure that the query is indeed catching true positive events
Baseline developmentDevelop a baseline (if needed) that will improve the detection fidelity
  • Baselines are sets of known and verified good behaviors and events present in the organization. Those events are normally excluded from the detection logic.
  • Baselines decisions and considerations should be documented and clearly stated in the ADS (Alerting and Detection Strategy)
  • Baselines are included in the hunt.yml/tf/hcl or alert.yml/tf/hcl files
Unittest DevelopmentThe unittest development is dependent on the type of devops pipeline. Simple goals are provided.Goals for the unittesting:
  • Changes or missing data
  • Syntax errors
  • To confirm detection logic by performing true positive detection
EnrichEnrich with additional data source if required
  • Each hunt could have different enrichment requirements. In some cases HR database could be used in order to understand if a person is on vacation, other trivial cases could be lookup of a hash, ip or domain in an threat intelligence repository etc.
Document
  • Create KB Document
  • Complete the ADS
  • MITRE ATT&CK coverage map update
  • Central knowledge base repository is required in order to mature the detection engineering program. This can be a github repository with controlled access that provides on a need to know basis the security teams members with access.
  • Each hunt should have a corresponding README.MD file that provides sufficient information and context. Consider an SOC analyst or Incident Responder responding to an event from your detection. By looking at the documentation they should be easily briefed on the premise and technicalities of the detection.
ValidateConfirm unittestsConfirm unittest are workingConfirmation of the unittests can be done by inspecting the implemented devops pipeline and ensuring that the actions (in the case of github) for unittests are running
True Positive validationValidate true positive event against real dataset using the query developed earlier.True positive validation can be achieved by:
  • Using historical event that exists in the central data repository
  • Emulation of the TTP by executing it in a controlled environment
False Positive ValidationEnsure no FP are produced by the query when ran against the prod dataset.
  • False positive events are good known events which are produced as output results by the detection/hunt query.
  • If baseline is used it should be validated that the baseline is catching those good known events. Splunk example: Splunk you can use makeresult command to create fake results and test your baseline and how you handle false positives.
AutomateAutomation & deploymentThis step is entirely dependent on the environment and should follow the standard ci/cd or automation practices of the organization.Integrate with devops pipeline and enable continuous deployment
ShareSocialize the new detectionA notification process is required and it should be created. The process can be in the form of newsletter or slack channel notification, preferably automated one.Follow a process to communicate the newly created detection with the Security Teams and inform them about it
Update Sec Dependency TreeThis document is actually part of the repository and can be shared with data engineering and security teams. The goal of sharing it is to promote care mentality where teams would check before they change. Meaning, if data engineer is about to rename an index they should first check if the index is being used. Having dependency document as part of the repository makes it easy and seamless for them to check.Update organization wide document showing dependencies for the detections

Process Flow

graph TD;
Research1(Opportunity Identification) -->Research2(Prioritize);
Research2 -->Research3(Develop Research Questions);
Research3 -->Research4(Information Gathering);
Research4 -->Research5(Collect Technical Context);
Research5 -->Prepare1(Identify Dataset);
Prepare1 -->Prepare2(Visibility Check);
Prepare2 -->Prepare3{Improve};
Prepare3 --> |yes| cis[Start security improvement initiative];
Prepare3 --> |no| Build1(Detection Query Creation);
Build1 --> Build2(Manual Testing);
Build2 --> Build3(Baseline development);
Build3 --> Build4(Automated Unittest Development);
Build4 -->Build5(Enrich);
Build5 --> Build6(Document);
Build6 -->  Validate1(Confirm unittests);
Validate1 -->val2(True/False Positive validation);
val2-->automate(Automation & deployment);
automate --> share(Socialize the new detection);
share -->share1(Update Sec Dependency Tree);

Midday

Phase 2️⃣ Midday ☀️

The “Midday” phase is normally the longest phase from the detection lifecycle, during which the detection has been engineered and commissioned to production. The phase monitors the detection during its operation and aims to improve it if needed. High level goals for the Midday phase:

  • Operate and monitor the detection for FP or TP
  • Improve the detection logic in case of influx of FP
  • Perform systematic reviews to ensure relevancy
FunctionsGoalDescriptionGuidelines
MonitorRun as per defined scheduleDetection is configured to run on pre-defined schedule or real time if applicableDetections will run based on the schedule set during the sunrise phase.
Confirm unittest passingMonitoring is configured to notify the responsible team in case the automation for the detection is not running properlySuggested approach: github actions - before deployment ensuring proper syntax
Work detectionsOnce detection is running it should be monitored for any TP or potential influx of FPTP events should be triaged, investigated and responded on by following an agreed IR process.
FP events should be investigated, proved as FP and documented as part of the baseline. Once the baseline is changed in the documentation the query can be updated and improved.
MeasureMeasure detection efficacyEnable metrics for the detection based on which areas for improvement can be identified.
Mitre Attack weakness
Success/failure of automating detections
Services covered
Each detection that covers particular TTP can be marked in the Mitre ATT&CK Navigator. Looking at percentage of covered tactics and techniques can be a metric.
Success or Failure in detection automation or influx of FP metric can be used to identify detections that require improvement.
Detection runtime length is a metric which can identify poorly written queries. For example, query too open that collects way too many events and chunks too much data only to spend even more time to filter by using custom logic.
Improve (optional)Improve detection fidelityOnce improvement opportunities have been identified during the operations or periodic review an improvement is triggeredThe goal of this function is to improve any detections which are with poor health (slow runtime, causing errors) and improve them by revisiting the detection logic.
ReviewPerform periodic reviewReview detections to identify improvement opportunities or decommission requirementsDetection can become irrelevant and thus decommissioned when:
The risk that it is compensating is far smaller than the cost of running the detection
The technology used for the detection is no longer present in the company

Midday phase Process Flow

graph TD;
Monitor1(Run per schedule) -->Monitor2(Receive and respond to alerts);
Monitor2 --> Monitor3{False Positives?} ;
Monitor3 --> |no| Measure[Document TP];
Measure --> Review(Perform periodic review)
Monitor3 --> |yes| Improve(Improve);
Improve --> Monitor1;

Sunset

Phase 3️⃣ Sunset 🌆

During the “Sunset” phase the detection is taken out of commission. The phase wants to ensure that resources are not spent for outdated detections that are no longer applicable and at the same time leave sufficient trace of the existence of the detection.

High level goals for the Sunset phase:

  • Decommission the detection and leave it in a state that it can be resumed anytime
  • Preserve knowledge
FunctionsGoalDescriptionGuidelines
DecommissionDecommission the detectionThe goal is to decommission the detection by following process that provides visibilityIn order to decommission a detection simply change the status field to "Sunset" in the .yml file. Assuming your devops pipeline is configured correctly, this should effectively disable the detections and prevent it from running.
Note: Do not remove anything from the repository as detections can be reused in future.
Knowledge base updateCreate an adequate indication in the KB document that the detection is no longer active and socialize the change with your security teams.Update Mitre coverage map by removing the coverage that the detection was providing

Sunset Process Flow

graph TD;
Review1[Review completed] --> Review2;
Review2{detection ready to decom} -->|no| End[end];
Review2{detection ready to decom} -->|yes| Preserve(Preserve knowledge);
Preserve --> Decommission(Decommission the detection);