Insights
Logging Strategy and Realistic Expectations

Two organisations can experience almost the same incident and be left with very different investigations. Both might discover that a privileged cloud account was compromised three months earlier. One recovers the administrative changes, data access, network modifications, and identity activity that followed. The other has a handful of sign-in records and little else because its control plane history expired, data access logging was never enabled, or relevant application logs were never collected.
Investigator skill can’t recover evidence that wasn’t collected or retained at the time of compromise. That’s why you have (or should have) a logging strategy. It determines which questions you’ll be able to answer later, how confidently you’ll be able to answer them, and which parts of the incident might remain unknowable. Your focus moves from collecting logs to understanding what an investigation will need. Instead of asking what logs should be collected, responders need to know what they must establish during an incident and whether available evidence supports those conclusions.
Logging Serves Several Different Jobs
You collect logs for detection, incident response, threat hunting, operational troubleshooting, auditing, policy enforcement, governance, and regulatory obligations. The uses overlap but aren’t interchangeable. The detection team might need a small number of high-value events delivered quickly enough to trigger an alert. An incident responder reconstructing activity months later might require a much broader evidence set with context that was unimportant when the event first occurred. An application owner might need verbose diagnostic output to troubleshoot a service, while an investigator needs stable identifiers, accurate timestamps, authentication context, and records of administrative changes.
NIST describes log management as the generation, transmission, storage, analysis, and disposal of log data. That places collection inside a larger lifecycle. The existence of an event at its source is only the first step. The event still has to be forwarded, retained, made available, protected, and understood before it can support an investigation. Compliance alone is therefore a poor measure of investigative readiness. An organisation might retain a required category of records for the required period and still be unable to answer how an account was used, what data was accessed, or which systems were affected. The records satisfy one purpose but not necessarily the one the response team now has.
A security platform might be great at detecting current activity while retaining too little history for a late-discovered compromise. Immediate detection and historical reconstruction are related capabilities, not the same capability. A logging strategy helps you decide which of these jobs matter, where they overlap, and where one requirement creates cost or compromise elsewhere.
Logs Only Answer Questions Within Their Collection Context
Logs record what source systems were configured to generate at that time, not everything that happened. Windows Event Logs and Sysmon can support findings about account use, process creation, service installation, PowerShell activity, network connections, and changes to system state. They only do that where relevant auditing, channels, and Sysmon rules were enabled before the activity occurred, though. Windows Event Forwarding (WEF) moves those records elsewhere but can’t create an event that the endpoint never generated. Microsoft describes WEF as passive: it doesn’t enable disabled channels, change audit policy, or enlarge local log files.
EDR adds process relationships, command lines, behavioural context, and estate-wide search. It’s often one of the most useful sources available during an investigation. It retains curated telemetry with a defined schema and provider-specific retention limits. Investigators must understand exactly which data points are preserved, for how long, and in what format. Microsoft Defender XDR retains data in the portal for 180 days, for example, while Advanced Hunting exposes 30 days of data unless it’s streamed elsewhere.
Identity records might establish that an account authenticated, the method used, the target resource, the apparent source address, and the result. They don’t automatically establish who was operating the account. Source addresses might represent proxies, secure access services, gateways, or other intermediaries. A successful authentication is evidence of account use within the logging system’s view. Attribution to a person requires more than that.
Network telemetry has similar limits. Firewall, DNS, proxy, and remote access logs might support findings about communication, name resolution, permitted connections, or access sequencing. They usually don’t prove which process initiated the traffic, what the user intended, or what occurred on the endpoint before and after the connection. Those conclusions need endpoint, identity, application, or cloud evidence alongside the network record.
Cloud logging makes the collection context especially visible. AWS CloudTrail event history provides 90 days of management events in each Region but doesn’t include data events or network activity events. Azure retains Activity Log events for 90 days and notes that the log typically doesn’t capture read operations. Google Cloud writes Admin Activity logs by default, while most Data Access logs are disabled unless explicitly enabled because their volume and cost can be substantial. An organisation can therefore have “cloud logging enabled” and still lack the records needed to determine which objects were read, what information was downloaded, or how data was used.
Empty Results Aren’t Simple Answers
An empty query result needs careful interpretation. It might mean that the activity didn’t occur. It might also mean that:
- The source never generated the event;
- The relevant category wasn’t enabled;
- The event remained local and wasn’t forwarded;
- Collection failed;
- The record was overwritten or expired;
- The event was stored somewhere else;
- The query used the wrong field or time range;
- Normalisation discarded useful detail;
- or the current interface doesn’t expose the retained record.
The gaps in the available record define what an investigator can reasonably conclude. A defensible finding states that no matching activity was identified in the available records for a particular period, explaining which records were available, their known coverage, and any material gaps. That’s more accurate than saying the activity never happened.
Collected logs aren’t necessarily complete either. Some sources write continuously, some emit events in batches, and others experience delivery delays. Systems might be offline. Legacy or OT devices might not support reliable forwarding. Ephemeral workloads might disappear before local evidence is collected. A central platform might remain available after an endpoint is destroyed, but it might hold only a subset of what existed locally. Timestamps require the same consideration. Millisecond precision can look authoritative, but precision isn’t the same thing as accuracy. System clocks drift. Services record different stages of a transaction. Events might be generated, queued, forwarded, normalised, and ingested at different times. Apparent sequence can support a conclusion without proving causality.
Logging Strategy is Capability Design
As ever, start with the questions. Which identities were used? Where did initial access occur? What executed, changed, communicated, or accessed data? How far did the activity spread, and did containment actually work? Those are the questions your evidence will have to answer. Then work backwards from there: Windows auditing, Sysmon, EDR, application records, and endpoint artefacts can each support parts of an execution timeline. Identity, endpoint, VPN, application, and cloud records can describe different parts of account use. Data access might require audit records that the default control plane history never contained.
The right mix depends on the environment. A cloud provider and a manufacturer don’t have the same systems, consequences, or collection options. The point is to know which sources can answer each question, where they can corroborate one another, and where the answer will still be incomplete.
Ownership is where I see a lot of logging programs often turn vague. The security team might run the central platform, but it doesn’t control every source. System owners configure local auditing. Application, identity, cloud, and infrastructure teams control what’s generated and whether it reaches a SIEM. Responders need authority to preserve it. Privacy, legal, and governance functions set other boundaries. Providers need clear obligations for records the organisation depends on.
A gap without an owner becomes an argument during the incident, while the evidence expires in the background. The record should say what’s missing and which incident question that prevents responders from answering. Give the gap an owner. Leadership can fund the fix or accept the limit, with a date to look at it again. That’s an operating decision, not a tuning detail.
The same idea applies across the lifecycle: generation, forwarding, retention, preservation, investigation, and review. Miss one stage and an investigative opportunity disappears. Detailed local events that’ve been overwritten can’t help you. Central retention doesn’t help if responders can’t retrieve or interpret the data. Searchable data hasn’t necessarily been preserved.
Logs need protection as well. They can contain sensitive identity, personal, and operational information, expose security architecture, and attract deletion or manipulation. Joint guidance from ASD, CISA, FBI, NSA, NCSC, and partners recommends central access, secure storage, integrity protection, restricted permissions, and separation from the systems being observed. Those measures help keep records available and reduce the risk of loss or tampering, but investigators still need to understand what was actually collected and retained.
More Telemetry Isn’t Automatically More Capability
“Collect everything” is unreasonable. It avoids prioritisation and doesn’t consider the constraints you’ll face. Storage and licensing costs rise, collection can affect performance, and analysts still have to find useful evidence in the noise. The cost also lands on endpoints, networks, privacy controls, schema maintenance, and analyst time.
Cloud data access logs, high-volume application telemetry, DNS analytics, and constrained systems all carry different costs. OT devices might not tolerate aggressive collection. Offline systems can require local retention and manual collection. Events excluded during generation or collection generally can’t be recovered later.
The same guidance recognises the forensic value of broad collection while recommending storage tiers based on how quickly data needs to be available. Keep detection data searchable. Move slower reconstruction data to cheaper storage. Where a source is too sensitive or too expensive to collect, document the decision and the investigative gap it creates.
The Incident Exposes What the Strategy Preserved
Consider a compromised Windows endpoint discovered four months after the initial activity. One organisation enabled relevant native auditing and Sysmon, forwarded selected events centrally, and streamed EDR telemetry to longer-term storage. Their responders can align process execution, account activity, service creation, and network connections. No source is flawless. Together, they can support a defensible reconstruction.
That capability wasn’t bought as one product. System owners enabled the sources, security retained the data, responders knew how to retrieve it, and leadership accepted the cost before the incident.
Another organisation relied mainly on its EDR hunting interface. The hunting window has passed, local Security logs have rolled over, and Sysmon was never deployed. Alerts or isolated artefacts might still exist, but the evidence is thinner. Claims about root cause or full scope have to reflect that. There are questions their record can’t answer.
The same applies to a suspicious network connection. DNS, firewall, and proxy records might establish that a host resolved and contacted a destination. On their own, they might not establish the initiating process or account. Endpoint and identity evidence can help connect the network activity to process execution and a particular session. Investigators should explain what the available evidence supports, identify any remaining gaps, and let their confidence reflect those limits.
Retention Isn’t Preservation
Retention keeps data available under a platform’s normal rules. Expiry, deletion, schema changes, broad access, or incomplete exports can still change what remains. Preservation starts when incident records need protection from those processes.
Good evidence handling means minimising changes, documenting collection, accounting for clock differences, retaining original format and metadata, verifying integrity, and restricting access so another examiner can understand how the evidence was handled. Legal requirements vary by jurisdiction. Incident plans should name who can suspend expiry and preserve local, cloud, or provider data, including who can approve longer retention. Waiting for those decisions during an incident costs time and might cost evidence.
Logging Capability Changes with the Environment
New services introduce different sources, schemas, and meanings. A documented retention setting doesn’t prove collection still works.
To test this in your organisation: pick an incident question and try to answer it with the data available today. Can we reconstruct simulated lateral movement? Does the cloud trail contain the data-access events required for critical assets? If the incident happened six months ago, do raw events remain to show who did what, whether remediation worked, and whether the relevant controls remain effective?
NCSC recommends doing this every 6 to 12 months, and whenever systems or their use change. NIST treats log management as an ongoing process. When a test finds a gap, give it an owner and make a decision: fund the change, require the provider to close it, or record why leadership accepts it. Then test it again when the environment changes.
Disclaimer: This content may have been edited or refined with assistance from AI tools. All final content, views, and recommendations are our own.