Useful security logging and monitoring comes down to a few decisions made in the right order. Collect logs from the sources that matter most: your identity provider, endpoints, cloud control plane, email and collaboration admin activity, network edge and critical applications. Make sure every system agrees on the time. Keep logs long enough to investigate something you discover months later. Then build a small number of well-tuned detections that someone actually responds to.
Most logging programs don't fail for lack of data. They fail because nobody can find the right event when it matters, or because alerts fire so often that people stop reading them. This guide is about avoiding both.
Why do so many logging projects disappoint?
The usual pattern goes like this: a team turns on everything, forwards it all to a central platform and waits for insight. The bill arrives before the insight does. Meanwhile the one log that would answer an investigator's question was never enabled, or it rolled over after seven days.
Visibility you can use starts from the questions you'll need to answer, not from the data you happen to have. Who signed in? From where? What did they change? What ran on this machine? What left the network? Each logging decision should serve one of those questions.
What should you log first?
If you're starting from scratch or rebuilding, prioritize sources by how much they tell you about access and change.
| Source | Key events to capture | Why it matters |
|---|---|---|
| Identity provider | Sign-ins, failed sign-ins, MFA registration and changes, admin role assignments, app consent grants, policy changes | Misuse of accounts, including stolen credentials, shows up here first |
| Endpoints | Process creation with command lines, logons, service and scheduled task creation, security tool status | Shows what actually executed and who ran it |
| Cloud control plane | API calls that create, modify or delete resources, IAM changes, logging configuration changes | Cloud changes happen through APIs, and misconfiguration is a common cause of exposure |
| Email and collaboration admin | Mailbox rules and forwarding, delegation, sharing settings, admin role changes, audit setting changes | These platforms hold much of your data and are administered through a web console |
| Network edge | Firewall allows and denies for exposed services, VPN authentication, DNS queries, web proxy | Shows inbound attempts against your perimeter and unusual outbound traffic |
| Critical applications | Admin actions, permission changes, bulk exports, authentication | Your finance, HR and customer systems are what an attacker ultimately wants |
Identity provider
If you can only do one thing, collect your identity provider's sign-in and audit logs. Include both interactive and non-interactive sign-ins where the platform separates them, and capture changes to MFA methods, conditional access policies and privileged roles.
Check your licensing too. Some platforms limit log retention or detail on lower tiers, so confirm what you're actually getting before you rely on it.
Endpoints
On Windows, make sure process creation auditing (event ID 4688) is enabled with command-line logging, along with logon events, account and group changes, service installation (7045) and audit log clearing (1102). On Linux, collect authentication logs and use auditd for process execution and changes to sensitive files. If you run an endpoint detection and response tool, its telemetry is usually richer than native logs, but make sure you can retain and search it for long enough.
Cloud control plane
Every major provider has an audit log for management activity: CloudTrail in AWS, the Activity Log in Azure and Cloud Audit Logs in Google Cloud. Turn it on for every account, subscription and project, in every region, and send it to a central location that ordinary administrators can't delete from. Data-plane logging, such as object-level access to storage, is valuable for sensitive buckets but gets expensive quickly, so enable it selectively.
Email and collaboration admin actions
Collaboration suites are often under-monitored because they feel like "just email". Admin logs here show mailbox forwarding rules to external addresses, delegated mailbox access, changes to external sharing settings, new admin roles and changes to audit settings. These are classic signs that a compromised account is being used to pull data out quietly.
Network edge
Firewall logs are noisy, so be selective. Capture traffic to and from internet-exposed services, VPN and remote access authentication, and DNS queries from internal resolvers. DNS logs in particular are compact and useful for tracing which machine talked to which domain.
Critical applications
Identify the handful of applications that hold your most sensitive data or money flows. Log administrative actions, permission changes and large exports. Many business applications have audit logging switched off by default.
Why does time synchronization matter?
An investigation is a timeline. If your firewall is four minutes ahead of your domain controller and your SaaS logs are in a different time zone, correlating events becomes guesswork.
The CIS Critical Security Controls (v8) call for at least two synchronized time sources across enterprise assets, where supported. In practice:
- Point servers, network devices and endpoints at a consistent, redundant set of NTP sources.
- Store timestamps in UTC in your central log platform, and convert to local time only for display.
- Know which timestamp each source records: event time, ingestion time or both.
- Alert on devices with significant clock drift.
How long should you keep logs?
Retention has to cover the gap between when something happens and when you find out. That gap can be months, especially with misuse of legitimate credentials.
A few reference points:
- CIS Controls v8 sets a minimum of 90 days for audit log retention.
- PCI DSS v4.0 requires at least 12 months of audit log history, with the most recent three months immediately available for analysis.
- OMB M-21-31, which applies to US federal agencies, calls for 12 months in active storage and 18 months in cold storage for many log types.
For most small and mid-sized organizations, a sensible target is 90 days searchable ("hot") and at least 12 months in cheaper archive storage. Keep identity and cloud audit logs at the longer end, since they're small and matter most in investigations. Check contracts, regulations and cyber insurance requirements, which may set their own minimums.
Protect the logs as well. Store them somewhere administrators of the monitored systems can't alter or delete, and alert when logging is disabled.
How do you turn logs into detections?
Collecting logs gives you the ability to investigate. Detection means something tells you to look. Start with a handful of high-value use cases rather than hundreds of generic rules.
Good starter use cases for most environments:
- A new privileged role assignment in the identity provider or cloud platform
- MFA disabled, or a new MFA method registered, for an administrator
- Many failed sign-ins followed by a success for the same account
- Audit logging disabled, a cloud audit trail stopped, or a Windows security log cleared
- A mailbox forwarding rule sending mail to an external domain
- Cloud storage made public, or a firewall rule opened to the whole internet
- Use of a cloud root or break-glass account
- Endpoint security tooling stopped or uninstalled
- Unusually large exports or downloads from a critical application
For each use case, write down the data source, the detection logic, who receives the alert and what they should do first. A detection without a response plan is just another log line.
Frameworks help here. MITRE ATT&CK gives you a shared vocabulary for mapping detections to attacker techniques and spotting gaps. The open-source Sigma project provides a large library of detection rules in a vendor-neutral format that you can adapt to your platform.
Test every detection. Generate the activity safely in a test account and confirm the alert fires, then retest after major platform changes.
How do you tune alerts so people keep reading them?
An alert queue that nobody trusts is worse than no queue at all, because it creates the impression that someone is watching. Tuning is ongoing work, not a one-time cleanup.
- Baseline first. Run new rules in a log-only mode for a week or two and see what normal looks like.
- Suppress with care. Exclude known-good activity narrowly, by specific account or host, and give exclusions an owner and a review date.
- Grade severity honestly. Reserve high severity for events that need someone to act now.
- Group related alerts. Ten alerts about the same user in five minutes should reach the analyst as one case.
- Track outcomes. Record whether each alert was a true positive, a benign true positive or a false positive. Rules that never produce anything useful should be fixed or retired.
- Write short runbooks. A few lines on what to check first makes after-hours response far more consistent.
How do you keep SIEM costs under control?
Most SIEM and log platforms charge by volume ingested, stored or searched. Costs creep up quietly as new sources are added.
- Find your top talkers. A small number of sources, often firewalls, DNS and verbose application logs, usually account for most of the volume.
- Filter at the source. Drop debug-level events and fields you never search. Aggregate repetitive allow traffic where you don't need every flow.
- Tier your storage. Send high-value security logs to the searchable tier and route bulky, rarely queried data to cheaper object storage you can search when needed.
- Avoid duplicates. The same event often arrives from two collectors or via two paths.
- Tie every source to a use case. If a log source doesn't support a detection, an investigation or a compliance requirement, question why you're paying to index it.
- Review monthly. Set a volume budget per source and alert on sudden jumps, which can also signal a misconfiguration.
Frequently asked questions
Do small organizations need a SIEM?
Not always a full one. What you need is centralized, searchable logs with enough retention and a few alerts that reach a person. That can be a lightweight log platform, the built-in alerting in your identity and cloud providers, or an outsourced monitoring service.
Should we log everything and filter later?
Rarely. Logging everything into an expensive search tier drives up cost without improving detection. Log broadly into cheap storage if you like, but be selective about what you index and alert on.
How often should someone review the logs?
The CIS Controls call for audit log reviews at least weekly. Your high-severity detections, though, need attention in hours, not days, which means someone has to own them outside business hours.
Next steps
- Confirm your identity provider, cloud audit logs and endpoint logs are flowing to one central, protected location.
- Check that every source uses synchronized time and that your platform stores timestamps in UTC.
- Set retention to at least 90 days searchable and 12 months archived, adjusting for your regulatory obligations.
- Pick five detection use cases from the list above, test them, and assign an owner and a runbook to each.
- Review alert outcomes and log volumes monthly, and retire what doesn't earn its keep.
Visibility isn't measured by how much data you store. It's measured by how quickly you can answer "what happened?" and how reliably someone notices when it matters.