Elevating Security Operations: The Critical Role of Data Quality in AI Performance
As artificial intelligence increasingly handles various workflows within security operations centers (SOCs), leaders in security find themselves asking a pivotal question: What has a greater impact on SOC performance—the advanced language model or the quality of the security data it processes?
Research in this area presents differing viewpoints. A study published in Frontiers in Artificial Intelligence highlights inconsistencies in the accuracy, relevance, and clarity of outputs from models like Claude 3.5, Gemini, and ChatGPT 4o. Conversely, findings in the Information Systems journal state that "trustworthy AI applications necessitate high-quality training and testing data across multiple quality dimensions, including accuracy, completeness, and consistency."
Recent research from the Provably Better Data project reinforces the data-centric argument. It demonstrates that the quality of data can be a decisive factor in the success of security operations, suggesting that high-fidelity network evidence can enhance security outcomes by 2 to 4 times.
Operational Benefits for CISOs
For Chief Information Security Officers (CISOs), the implications of these findings are clear and actionable:
- Concrete telemetry improves the mean time to respond (MTTR).
- Better data control can help manage token expenditures, leading to cost efficiency.
- Enhanced return on security investments provides proof of security team's effectiveness.
Security professionals should evaluate these insights to inform architectural decisions in their SOC. Armed with this data, CISOs can better justify infrastructure spends, alleviate analyst burnout from alert fatigue, and present quantifiable security metrics to executive teams and boards.
Testing the Hypothesis
To identify the core drivers of AI effectiveness in security operations, the Provably Better Data project devised a controlled testing framework.
The study employed two main evaluation benchmarks:
- A Capture the Flag (CTF) scenario featuring a 44-question investigation based on a Volt Typhoon attack campaign.
- An incident response analysis requiring report generation from a Salt Typhoon dataset.
To isolate data quality as the defining variable, the project assessed four distinct telemetry types under uniform conditions:
- Corelight enriched logs
- Open-source nDPI firewall logs
- Snort 3 intrusion detection system alerts
- NetFlow connection telemetry
All datasets were evaluated multiple times per an Open Cybersecurity Schema Framework (OCSF) normalized schema. The LLMs tested included Anthropic Claude Opus 4.6, Google Gemini Pro 3.1 Preview, and older versions of both models, assessed equally across every test backdrop.
Insights: Limits of Evidence in Automated Processes
While modern language models excel in reasoning, their conclusions are ultimately constrained by the quality of the data available. Missing crucial protocol-level context means that AI cannot derive insights that haven't been captured. The depth of investigation is limited by the evidence at hand.
Consider an example from the CTF to illustrate this point. When querying for the NetBIOS computer name associated with IP address 10.110.154.113:
- Corelight logs successfully returned FINANCE01, derived from the server_nb_computer_name field within Corelight’s NTLM log.
- Firewall logs recognized NTLM activity but couldn't extract specific fields, failing to provide the correct name.
These data deficiencies can introduce friction in operations. While firewall logs contribute to broad connection visibility, the absence of detailed protocol information hampers effectiveness. Thus, when available evidence is lacking, analysts must manually validate AI outputs. In contrast, sufficient data enables immediate actionable insights from AI.
Measurable Improvements: Data Quality Counts
The research displayed that superior data quality correlates directly with significantly enhanced security outcomes, achieving improvements of 2 to 4 times in effectiveness.
During the 44-question CTF assessment, accuracy rates varied considerably based on data sources:
- Corelight logs exhibited a 95.2% accuracy rate.
- Firewall logs yielded a 58.3% accuracy rate.
- Snort 3 alerts documented a 39.4% accuracy rate.
- NetFlow records reported only a 25.8% accuracy rate.
Access to richer data directly influenced model performance. With Corelight data, each language model could respond accurately to all 44 questions, whereas NetFlow logs only enabled responses to 15 questions. Consequently, Corelight provided a CTF score of 4,178.3 points, greatly surpassing the 970.0 points of NetFlow.
The results in the incident response experiment reflected similar disparities across evidence coverage and significant findings:
- Corelight logs documented a 90.3% evidence coverage rate.
- Firewall logs had a 61.3% coverage rate.
- NetFlow records displayed a coverage rate of 30.9%.
- Snort 3 alerts reached only a 21.2% coverage rate.
For critical Tier 1 investigation requirements, Corelight logs supported answers to 91.7% of mandatory questions, whereas NetFlow logs covered only 18.3% and Snort 3 alerts merely 10%.
The outcome: High-quality data produced a fivefold increase in critical incident visibility compared to basic flow metrics.
Investigation speed was also notably faster with enriched data. The LLM concluded full investigations in 14.7 minutes using Corelight logs, while it took 27.0 minutes for NetFlow and 26.3 minutes for firewall logs.
The outcome: Subpar data almost doubled investigation durations due to repeated retry loops in models.
Interestingly, models rarely generated erroneous outputs when directed to mark missing evidence as unanswerable. The counts of hallucinations were zero for Corelight, firewall, and NetFlow datasets, while Snort 3 alerts reported 1.9 instances. Grounded models help analysts avoid pursuing fictitious evidence, thereby conserving investigation times and enhancing confidence.
Strategic Recommendations for Security Leaders
Automation through AI can expedite threat detection, indicator monitoring, and incident response across networks. However, achieving effective automation doesn’t happen by chance; it demands comprehensive, structured, and protocol-aware telemetry to yield dependable results.
The research implies that merely upgrading models and refining prompts won’t rectify fundamental deficiencies in data quality. Therefore, SOC leaders contemplating future investments should prioritize evidence quality over model selection to enhance operations significantly.
For a more detailed exploration, including methodologies and complete metrics, consult the Provably Better Data white paper on Corelight’s website.
Corelight Network Detection and Response
Data quality can elevate security outcomes significantly. Understand why high-fidelity network evidence is essential for AI-driven security strategies. Corelight: safeguarding the world's most sensitive networks. Discover more.