A Stage-Aware Tool for Modelling and Verifying CI/CD Security Evidence

Sabbir M. Saleh, Computer Science, University of Western Ontario, London, ON, Canada, ssaleh47@uwo.ca
Nazim H. Madhavji, Computer Science, University of Western Ontario, London, ON, Canada, madhavji@gmail.com
John Steinbacher, Cloud Division, IBM Canada Lab, Markham, ON, Canada, jstein@ca.ibm.com

Continuous integration and continuous deployment pipelines produce security-relevant records across source, build, test, package, deployment, and runtime environments. However, these records remain fragmented across logs, scanner findings, policy decisions, identity events, artefact metadata, and runtime alerts, limiting their usefulness during investigation. We present DevSecLogs, a stage-aware tool for modelling and verifying CI/CD security evidence. DevSecLogs normalises heterogeneous pipeline records into structured evidence objects that connect an observation to its pipeline stage, originating artefact or activity, security requirement, deterministic reason code, priority level, and verification status. The tool combines semantic triage and anomaly-oriented analysis with evidence-checkable explanations rather than presenting uninterpreted model outputs. Selected evidence bundles are hashed for consistency checking; the browser demonstration uses local anchors, while external anchoring remains experimental. Through an analyst-facing interface, users inspect runs, cross-stage findings, supporting evidence, and verification results. The demonstration illustrates an end-to-end investigation involving a skipped test gate, premature artefact publication, and an artefact-integrity inconsistency. DevSecLogs operationalises a requirements-to-evidence model that complements existing CI/CD scanners, policy engines, signing systems, and monitoring platforms.

CCS Concepts: • Security and privacy → Software security engineering; • Software and its engineering → Software development process management;

Keywords: CI/CD security, evidence modelling, DevSecOps, software supply chain, security investigation, evidence integrity

ACM Reference Format:
Sabbir M. Saleh, Nazim H. Madhavji, and John Steinbacher. 2026. A Stage-Aware Tool for Modelling and Verifying CI/CD Security Evidence. In ACM/IEEE 29th International Conference on Model Driven Engineering Languages and Systems (MODELS Companion 2026), October 04--09, 2026, Málaga, Spain. ACM, New York, NY, USA 5 Pages. https://doi.org/10.1145/3837062.3838889

1 Introduction

Continuous integration and continuous deployment (CI/CD) pipelines coordinate source-code changes, dependency retrieval, compilation, testing, artefact packaging, deployment, and runtime operation [1]. Each activity produces security-relevant information, including workflow events, scanner findings, test results, identity records, artefact metadata, policy decisions, and runtime alerts. Although these records may reveal insecure behaviour, they are commonly distributed across heterogeneous tools and storage systems, a concern reflected in CI/CD guidance and large-scale analyses of workflow overprivilege and third-party action use [2, 5]. Consequently, an investigator may know that a scanner reported a vulnerability, a test was skipped, or an artefact digest changed, but still lack an explicit representation of where the observation occurred, which requirement it affects, how it relates to other pipeline events, and whether the supporting evidence has remained unchanged.

Existing CI/CD security mechanisms address important parts of this problem. Scanners identify vulnerable dependencies or insecure code, policy engines enforce deployment conditions, signing systems establish artefact provenance, and observability platforms aggregate logs and alerts. However, their outputs often remain isolated and tool-specific. A scanner finding does not inherently describe its relationship to a subsequent deployment, while an anomaly score does not explain which pipeline obligation was violated. Similarly, preserving a log or its digest does not make the record investigation-ready unless its stage, affected asset, security meaning, and relationship to other evidence are represented explicitly. Supply-chain mechanisms such as Sigstore provide complementary signing and transparency capabilities [4]. However, they do not by themselves organise heterogeneous scanner, policy, identity, and runtime records into a common investigation model.

This paper presents DevSecLogs, a research prototype that operationalises a stage-aware evidence model connecting observations to run, lifecycle stage, source, actor, asset, requirement, supporting evidence, deterministic reason code, priority, and verification status. It normalises heterogeneous records across source, build, test, package, deploy, and runtime, then supports AI-assisted triage, deterministic explanation, cross-stage linking, and integrity verification. DevSecLogs is an investigation layer, not a replacement for scanners, policy engines, signing tools, or security information and event management systems.

Our earlier work introduced DevSecLogs as a CI/CD log-intelligence vision [6]; separate studies examined topic modelling [7], sequential anomaly detection [10], and blockchain-supported evidence management [8]. Relative to that work and the requirements-to-evidence framework [9], this paper contributes an implemented evidence metamodel, executable reason-code evaluation, cross-stage linking, an analyst-facing workflow, and post-capture verification. AI outputs remain prioritisation signals; investigation-ready findings require observable evidence, pipeline context, and inspectable reasons.

The modelling contribution lies in explicit relationships among activities, requirements, observations, artefacts, reasons, and verification records. Publication before completion of a required test maps to PUBLISH_BEFORE_TEST, while a digest mismatch maps to CHECKSUM_MISMATCH. These are defined evidence conditions rather than free-form explanations, and related objects can be grouped into cross-stage bundles for integrity checking.

This paper makes three contributions:

  1. a stage-aware CI/CD security-evidence model connecting heterogeneous observations with lifecycle stages, actors, assets, requirements, reason codes, priorities, and verification metadata;
  2. an implementation of the model in an analyst-facing prototype supporting evidence ingestion, stage normalisation, AI-assisted prioritisation, deterministic explanation, cross-stage linking, and integrity verification; and
  3. an end-to-end demonstration and preliminary validation using representative CI/CD security scenarios.

2 Stage-Aware Security Evidence Model

DevSecLogs represents CI/CD security information through a stage-aware evidence model. The model provides a common structure for records originating from repository events, build logs, test outcomes, scanner reports, registry actions, deployment policies, identity traces, and runtime telemetry. It retains the contextual and semantic properties needed to determine where an observation occurred, what it concerns, why it is security-relevant, and whether its recorded state can later be verified. The metamodel is domain-specific, but its separation of activities, agents, entities, relations, and bundles is informed by W3C PROV-DM [3]; its explicit connection between requirements and supporting evidence is also related to SACM-based assurance modelling [12]. We use these as conceptual neighbours and do not claim formal conformance.

Figure 1
Figure 1: Stage-aware security-evidence metamodel linking observations to pipeline context, requirements, reasons, bundles, and verification records.

Metamodel versus serialisation. The metamodel and the JSON representation serve different purposes. JSON is a concrete serialisation used by the prototype, and a JSON schema can validate field presence, types, and enumerated values. The metamodel defines the domain semantics that span records: class relationships and multiplicities, stage membership, requirement-to-observation links, bundle membership, and temporal or shared-asset relations. For example, every EvidenceObject belongs to one pipeline run and one primary stage; PUBLISH_BEFORE_TEST requires an explicit precedence relation between publication and a required test; and a BundleRelation must be justified by a shared run, asset, actor, requirement, or temporal dependency. These investigation constraints are not provided by the serialisation format alone.

Pipeline context. A pipeline run R is associated with activities across six logical lifecycle stages:

Math 1
These stages provide a stable abstraction over platform-specific job and workflow names. The model does not assume that all stages execute once or in a strictly linear order. Pipeline jobs may be repeated, omitted, branched, or executed concurrently. DevSecLogs therefore records both the logical stage and observed event time, while temporal reason codes are evaluated using precedence relations among specific events.

Evidence object. The principal unit of analysis is an evidence object:

Math 2
Here, run identifies the pipeline execution; stage provides lifecycle context; source identifies the producing system or record type; actor represents a user, service account, token, or automated process; and asset identifies the affected repository, dependency, artefact, image, deployment, or runtime resource. The requirement field references the security obligation against which the observation is interpreted. The observation preserves the relevant record or evidence line, while reason explains why that observation warrants investigation. The priority value supports triage but is not treated as a calibrated incident-risk estimate. The final fields describe the evidence commitment and its current verification state. In the demonstration build, priority is a transparent triage heuristic: a clamped weighted sum of matched reason codes, while Low, Medium, High, and Critical labels are assigned by configured reason combinations. It is not presented as a calibrated probability of compromise.

A raw record becomes an evidence object only after sufficient context has been assigned. At minimum, an object identifies its pipeline run, primary stage, source, observation, and security interpretation. Missing contextual properties are represented explicitly as unknown rather than inferred without evidence.

Requirements and reason codes. A requirement represents an expected security property, such as “an artefact must not be published before required tests complete”, “a deployed image must possess an approved signature”, or “a pipeline identity must use only the permissions required for its task”. Requirements may originate from pipeline policy, organisational controls, scenario-specific security specifications, or recognised CI/CD guidance such as NIST SP 800-204D and the OWASP CI/CD risk taxonomy [1, 5].

A reason code denotes an inspectable condition relating an observation to a requirement. For example,

Math 3
where p is a package-publication event and t is a required test activity. Similarly, CHECKSUM_MISMATCH is assigned when an observed artefact digest differs from the expected digest under the applicable integrity requirement. Other supported reason categories cover failed signature verification, policy rejection of unsigned images, excessive pipeline privileges, and evidence-verification failure.

Reason codes are deterministic and evidence-checkable. Topic labels and anomaly scores may identify records requiring attention, but they do not independently determine the reason code. A reason is assigned only when its defined observable condition is satisfied.

Cross-stage evidence bundles. Individual evidence objects may be related through a shared run, actor, asset, requirement, or temporal dependency. DevSecLogs groups relevant objects into an evidence bundle:

Math 4
where EB is a non-empty set of evidence objects and relations records why they are connected. A bundle may associate a skipped test, premature package publication, failed artefact verification, and subsequent deployment event. Such a bundle represents an investigation pathway, not a confirmed attack.

For selected run bundles, the implementation sorts normalised events by stage order, file name, and line number; projects each event to its stage, timestamp, message, and reason codes; serialises the resulting array as JSON; and computes a SHA-256 digest. Recomputing the digest supports post-capture consistency checking. The current browser demonstration stores both evidence state and its anchor locally; therefore, it detects accidental or non-coordinated modification but does not resist an adversary who can modify both values. Strong adversarial tamper evidence requires an independently protected commitment, such as the experimental Hyperledger Fabric path [8]. Verification also does not prove that the originating record was complete or truthful before capture.

3 Tool Architecture and Workflow

DevSecLogs implements the evidence model through a modular architecture comprising an analyst interface, evidence-processing services, AI-assisted analysis, requirements reasoning, evidence management, and cloud deployment. Figure 2 presents the six-step workflow and a representative prototype alert view that exposes the run and stage context, deterministic reason codes, digest evidence, and verification result.

Figure 2
Figure 2: DevSecLogs implementation: (a) evidence-processing workflow; (b) alert view and verification.

Architecture. The browser-based interface provides dashboard, log-analysis, alert, run, and evidence-ledger views. A separately deployed Flask service supports topic-model inference, while the demonstration application performs stage mapping, deterministic reason evaluation, bundle construction, and digest verification. This separation prevents probabilistic topic or anomaly signals from independently deciding why an observation constitutes security evidence.

Model realisation. Evidence objects and run bundles are represented as JSON records, but the JSON representation is an implementation carrier rather than the model contribution itself. Source-specific rules instantiate the stage, actor, and asset associations; auditable predicates and regular-expression rules instantiate security requirements and reason codes; and linking rules create explicit cross-object relations. The demonstration recomputes a SHA-256 hash chain and reports PASS or FAIL. The browser-local anchor provides a consistency check, not adversarially independent protection; the Hyperledger Fabric integration remains the experimental external-anchoring path [8].

The evidence layer complements supply-chain attestation and signing systems. For example, in-toto link metadata and Sigstore verification results can enter the model as observations alongside test, policy, identity, scanner, and runtime records [4, 11].

Running example. The demonstration uses prepared records for pipeline tekton/devseclogs, run run-102, and container artefact api:42. The fixture contains a required test still running when the artefact is published, an expected/observed digest mismatch, and an outbound destination outside the configured allowlist. This single run is used throughout the interface to show how raw observations become model instances and then a linked investigation bundle.

Evidence-processing workflow. Collect accepts pasted log excerpts, structured records, and preloaded stage-specific samples. Each input is associated with available metadata, including pipeline, run identifier, source type, timestamp, and originating platform.

During Normalise, platform-specific records are mapped to the common JSON evidence schema. Stage inference uses configurable filename and content rules, while available actors, assets, event types, timestamps, and evidence lines are retained. Unknown properties remain unspecified rather than being inferred without supporting evidence.

Analyse generates prioritisation signals. Topic modelling groups textual records into semantically related security themes, while anomaly-oriented processing identifies behaviour that differs from expected patterns. These outputs are supporting attributes rather than final security decisions.

During Explain, the tool evaluates evidence conditions associated with security requirements. When a condition is satisfied, the corresponding deterministic reason code is attached to the evidence object. The interface displays the reason code with its requirement, affected stage, evidence line, and priority level so that an analyst can inspect the basis of the finding.

Link connects evidence objects using shared pipeline runs, artefacts, actors, requirements, and temporal relationships. A skipped test, premature package publication, failed signature check, and policy-denied deployment may therefore be presented as a connected investigation pathway rather than independent alerts. The link does not by itself establish that an attack occurred.

Finally, Verify recomputes the canonical bundle and SHA-256 hash chain. The demonstration reports PASS when the recomputed last hash matches the stored anchor and FAIL otherwise.

4 Demonstration and Preliminary Validation

The demonstration addresses the following evaluation question:

EQ1: Can DevSecLogs construct, link, explain, and verify stage-aware security evidence for representative CI/CD security scenarios?

The primary scenario replays the prepared run-102 fixture introduced in Section 3. A test record at 11:35:09 remains active when api:42 is published at 11:35:12; DevSecLogs therefore links the two observations and assigns PUBLISH_BEFORE_TEST. A subsequent record contains different expected and observed SHA-256 digests, producing CHECKSUM_MISMATCH; an outbound destination outside the fixture allowlist produces EGRESS_NOT_ALLOWLISTED. The tool connects these observations through the shared run, artefact, stages, and temporal relations. The analyst then inspects the bundle and performs the browser-local consistency check, obtaining PASS before a controlled modification and FAIL afterwards.

Table 1: Controlled demonstration fixtures and observed outcomes.
Case Expected interpretation Observed behaviour
Premature publication Artefact published before a required test completed Mapped test and publication records and assigned PUBLISH_BEFORE_TEST.
Digest inconsistency Observed digest differs from expected digest Linked package and deployment evidence and assigned CHECKSUM_MISMATCH.
Unsigned deployment Policy rejects an image without an approved signature Assigned COSIGN_VERIFY_FAIL and OPA_DENY_UNSIGNED_IMAGE.
Suspicious egress Outbound connection is not allowlisted Assigned EGRESS_NOT_ALLOWLISTED to the runtime/build evidence.
Evidence modification Locally anchored evidence changes after capture Recomputed the hash chain and changed the local consistency result from PASS to FAIL.

4.1 Model-to-Tool Traceability

Table 2 shows how principal metamodel constructs are realised and exercised rather than serving only as a descriptive diagram.

Table 2: Metamodel-to-tool traceability.
Construct Prototype realisation Demonstrated behaviour
PipelineStage Stage-mapping rules and common schema Maps records to source, build, test, package, deploy, or runtime.
EvidenceObject Structured JSON evidence record Preserves run, stage, actor, asset, observation, reason, digest, and status.
Requirement Auditable requirement predicates Associates findings with test, integrity, signature, or network obligations.
ReasonCode Deterministic rule evaluation Produces inspectable codes such as PUBLISH_BEFORE_TEST.
BundleRelation Links by run, asset, actor, stage, and time Connects otherwise separate cross-stage evidence.
EvidenceBundle Stored evidence and relations Presents an investigation pathway without asserting an attack.
VerificationRecord Canonicalisation, SHA-256, and comparison Returns PASS for unchanged and FAIL for modified local evidence.

4.2 Matched Non-Triggering Controls

Matched controls complete testing before publication, match expected and observed digests, use an allowlisted egress destination, and validate an approved signature. In each case, the corresponding reason code should remain absent. These controls exercise rule specificity at the fixture level; they are not a statistical estimate of operational false-positive rates.

4.3 Validation Scope

The fixtures test functional conformance between the metamodel, rule implementation, and expected tool behaviour. Triggering cases exercise rule firing, while matched controls define when reasons must remain absent. The local SHA-256 check covers non-coordinated modification only; an adversary able to change both evidence and its browser-local anchor is outside this guarantee. Live connectors and independently protected commitments remain partial or experimental. The evaluation therefore does not establish statistical detection accuracy, analyst-efficiency gains, adversarial tamper resistance, or production readiness.

5 Discussion and Conclusion

DevSecLogs complements scanners, policy engines, provenance systems, and observability platforms by connecting heterogeneous security outputs to runs, stages, assets, requirements, deterministic reasons, bundles, and verification records. The demonstration shows model instantiation, cross-stage linking, rule evaluation, bundle construction, and local consistency checking while keeping prioritisation signals separate from evidence-checkable security interpretation. Adversarial tamper evidence requires independently protected commitments. Future work will harden external anchoring, add native connectors, improve concurrent-event correlation, and evaluate investigation usefulness with practitioners.

References

  • Ramaswamy Chandramouli, Frederick Kautz, and Santiago Torres-Arias. 2024. Strategies for the Integration of Software Supply Chain Security in DevSecOps CI/CD Pipelines. Technical Report NIST Special Publication 800-204D. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-204D
  • Igibek Koishybayev, Aleksandr Nahapetyan, Raima Zachariah, Siddharth Muralee, Bradley Reaves, Alexandros Kapravelos, and Aravind Machiry. 2022. Characterizing the Security of GitHub CI Workflows. In 31st USENIX Security Symposium (USENIX Security 22). USENIX Association, 2747–2763. https://www.usenix.org/conference/usenixsecurity22/presentation/koishybayev
  • Luc Moreau and Paolo Missier. 2013. PROV-DM: The PROV Data Model. W3C Recommendation. World Wide Web Consortium. https://www.w3.org/TR/prov-dm/
  • Zachary Newman, John Speed Meyers, and Santiago Torres-Arias. 2022. Sigstore: Software Signing for Everybody. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. 2353–2367. https://doi.org/10.1145/3548606.3560596
  • OWASP Foundation. 2022. OWASP Top 10 CI/CD Security Risks. OWASP Project, version 1.0. Accessed 13 July 2026. https://owasp.org/www-project-top-10-ci-cd-security-risks/
  • Sabbir M. Saleh. 2025. DevSecLogs: AI-Powered, Tamper-Evident Log Intelligence for Secure CI/CD Pipelines. In IEEE International Conference on Software Maintenance and Evolution (ICSME). 890–894. https://doi.org/10.1109/ICSME64153.2025.00101
  • Sabbir M. Saleh, Nazim H. Madhavji, and John Steinbacher. 2024. Enhancing Cloud Security Through Topic Modelling. In 28th ACIS International Winter Conference on Software Engineering, Artificial Intelligence, Networking and Parallel/Distributed Computing (SNPD-Winter). 65–71. https://doi.org/10.1109/SNPD-Winter67015.2024.00020
  • Sabbir M. Saleh, Nazim H. Madhavji, and John Steinbacher. 2025. Towards a Blockchain-Based CI/CD Framework to Enhance Security in Cloud Environments. In 20th International Conference on Evaluation of Novel Approaches to Software Engineering (ENASE 2025). https://doi.org/10.5220/0013298200003928
  • Sabbir M. Saleh, Nazim H. Madhavji, and John Steinbacher. 2026. A Stage-Aware Requirements-to-Evidence Framework for Secure CI/CD Pipelines. In 13th International Workshop on Evolving Security and Privacy Requirements Engineering (ESPRE ’26). To appear.
  • Sabbir M. Saleh, Ibrahim Mohammed Sayem, Nazim H. Madhavji, and John Steinbacher. 2024. Advancing Software Security and Reliability in Cloud Platforms through AI-Based Anomaly Detection. In Cloud Computing Security Workshop (CCSW ’24). https://doi.org/10.1145/3689938.3694779
  • Santiago Torres-Arias, Hammad Afzali, Trishank Karthik Kuppusamy, Reza Curtmola, and Justin Cappos. 2019. in-toto: Providing Farm-to-Table Guarantees for Bits and Bytes. In 28th USENIX Security Symposium (USENIX Security 19). USENIX Association, 1393–1410.
  • Ran Wei, Tim P. Kelly, Xiaotian Dai, Shuai Zhao, and Richard Hawkins. 2019. Model Based System Assurance Using the Structured Assurance Case Metamodel. Journal of Systems and Software 154 (2019), 211–233. https://doi.org/10.1016/j.jss.2019.05.013

CC-BY license image
This work is licensed under a Creative Commons Attribution 4.0 International License.

MODELS Companion 2026, Málaga, Spain

© 2026 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-2903-4/26/10.
DOI: https://doi.org/10.1145/3837062.3838889