Knowledge Hub

Secure Data Pipelines: Why Most Breaches Come Through Credentials, Not Encryption Gaps

Secure Data Pipelines: Why Most Breaches Come Through Credentials, Not Encryption Gaps

Table of Contents

In 2024, attackers stole data from Snowflake customer environments without breaking any encryption and without breaching Snowflake itself. They logged in. Mandiant traced the campaign to credentials harvested by infostealer malware, used against accounts that had no multi-factor authentication, and Snowflake and Mandiant notified roughly 165 potentially exposed customers. That is the pattern this guide is built around. Most advice on securing data pipelines starts with encryption. Most recent pipeline breaches started with an identity, a leaked secret, or a compromised piece of pipeline tooling. This guide covers what a secure data pipeline actually requires, where pipelines really get breached, the controls that matter at each stage, and how to audit a pipeline you already run.

What is a secure data pipeline?

A secure data pipeline is one where every component that moves, transforms, or stores data, and every identity and piece of code that runs it, is controlled so that only authorized systems and people can read, change, or redirect the data. Encryption protects the data itself. Pipeline security also has to protect the credentials, service accounts, scripts, and third-party tools that have access to it, because those are where most attacks begin.

One clarification, because the terms get confused in search results: a secure data pipeline is any business data pipeline built with security controls. A security data pipeline is a specific category of tooling that routes security logs and telemetry into monitoring platforms such as a SIEM. This guide is about the first: protecting the pipelines that move your operational, customer, and analytics data.

Where do data pipelines actually get breached?

Data pipelines are most often breached through three routes: stolen or weak credentials on the systems the pipeline connects to, secrets such as API keys and passwords left in code or configuration, and compromised third-party tools in the build and deployment process. All three bypass encryption entirely, because the attacker arrives with legitimate access.

  • Identity: the Snowflake campaign (2024). The threat group Mandiant tracks as UNC5537 used stolen credentials to access more than 100 Snowflake customer tenants. The affected accounts lacked multi-factor authentication, the credentials were still valid, and the instances had no network policies restricting access to trusted locations. Snowflake’s own systems were not breached.
  • Secrets in code: secrets sprawl. GitGuardian’s 2025 State of Secrets Sprawl report found nearly 24 million new hardcoded secrets exposed on public GitHub in 2024, a 25% increase on the year before, and that 70% of secrets leaked in 2022 were still valid. Pipeline code, connection strings, and orchestration configs are exactly where these secrets live.
  • Pipeline tooling: the tj-actions compromise (2025). In March 2025, the widely used GitHub Action tj-actions/changed-files, used across more than 23,000 repositories, was compromised to extract secrets from continuous integration runs, including access keys and tokens, and CISA issued an alert. Any data pipeline deployed through an affected workflow was exposed through its build process, not its data.
SECURE DATA PIPELINES · THE EXPOSURE

Four Numbers on Where Pipelines Really Break

The costly incidents rarely involve cracking encryption. They involve access someone should not have had.

$4.44M

Average breach cost

Global, 2025; USD 10.22M in the US

~165

Customers notified

Snowflake campaign, stolen credentials, no MFA

~24M

Secrets leaked

On public GitHub in 2024 alone

23,000+

Repositories exposed

One compromised CI action, March 2025

Source: IBM Cost of a Data Breach Report (2025); Mandiant and Snowflake (2024); GitGuardian State of Secrets Sprawl (2025); CISA alert on CVE-2025-30066 (March 2025).

The cost of getting this wrong is not abstract. IBM’s 2025 Cost of a Data Breach Report puts the global average cost of a breach at USD 4.44 million, and the US average at a record USD 10.22 million.

What security controls does each stage of a data pipeline need?

Each stage of a data pipeline carries a different risk, so the controls differ by stage: ingestion needs authenticated, validated sources; transport needs encryption; transformation and storage need least-privilege access; consumption needs governed sharing; and the orchestration and deployment layer needs secrets management and dependency control. The last row is the one most security checklists skip.

Stage Main risk Controls that matter
Ingestion Untrusted or spoofed sources, malformed data Authenticated source connections, input validation, schema checks
Transport Interception in transit TLS for every hop, private network paths where possible
Transformation Over-privileged jobs that can read or alter everything Least-privilege service accounts scoped per job
Storage Account takeover, unrestricted access to the warehouse MFA, network policies limiting access to trusted locations, encryption at rest, masking of sensitive fields
Consumption Over-shared extracts and dashboards Role-based access, row-level security, logged exports
Orchestration and deployment Leaked secrets, compromised dependencies and CI tools A secrets manager, pinned dependency versions, isolated build runners, audit logging

Encryption in transit and at rest is table stakes and should be on by default everywhere. It is also the control that would not have stopped any of the three incidents above.

How do you secure the identities that run a data pipeline?

Secure pipeline identities by giving every pipeline component its own narrowly scoped account, enforcing multi-factor authentication for every human login to data platforms, restricting where accounts can connect from, and rotating credentials on a schedule. The Snowflake campaign succeeded because each of these was missing on the affected accounts.

  • One service account per job, scoped to what it needs. A transformation job that only reads two tables should not hold credentials that can read the entire warehouse. Shared, all-powerful service accounts turn one leaked credential into full access.
  • MFA on every human account. Analysts, engineers, and administrators who log in to the warehouse or orchestration tool should not be able to do so with a password alone.
  • Network policies. Restrict which IP ranges or private networks can reach the data platform. A valid credential used from an unexpected location should fail.
  • Rotation and offboarding. Credentials that never expire stay valid after the person or system that held them is gone. GitGuardian’s finding that 70% of secrets leaked in 2022 were still valid two years later shows how rarely this happens in practice.

Access tends to accumulate over time, which is the same over-permissioning problem we cover in our guide to data accessibility. A quarterly review of which accounts can reach which data is the cheapest control on this list.

How do you secure the code and tools a pipeline depends on?

Secure pipeline code by keeping secrets out of it entirely, pinning every third-party dependency and CI action to a verified version, isolating build environments, and scanning repositories for leaked credentials before and after every commit. Pipelines are software, and they inherit every risk of the software supply chain.

  1. Move every secret into a secrets manager. Connection strings, API keys, and passwords belong in a dedicated secrets store that the pipeline reads at runtime, never in code, notebooks, or configuration files committed to a repository.
  2. Pin dependencies and CI actions to exact versions. The tj-actions compromise worked because many workflows referenced the action by a movable tag. Pinning to a specific commit means a malicious update cannot silently change what runs in your pipeline.
  3. Isolate and minimize build runners. Build and deployment jobs should hold only the credentials they need for that run, and should not be able to reach production data unless the job genuinely requires it.
  4. Scan for secrets continuously. Automated secret scanning on every commit, plus a periodic scan of repository history, catches credentials before an attacker does. Any secret found should be treated as compromised and rotated, not just deleted.

These controls also make pipelines easier to maintain. A pipeline whose credentials and dependencies are explicit and versioned is easier to debug when it breaks, which ties directly to the maintenance problems covered in our guide to data transformation.

How do you audit the security of a data pipeline you already run?

Audit an existing pipeline by mapping every system and data source it touches, listing every identity and secret it uses, checking each against the controls above, and fixing the gaps in order of exposure. Start with what an attacker would try first: human logins without MFA, credentials in code, and over-privileged service accounts.

  1. Map the pipeline end to end. List every source, transformation job, storage layer, and consumer, including the ones nobody officially owns. Unknown components cannot be secured, which is why a data discovery pass is often the first step.
  2. Inventory identities and secrets. For each component, record which account it runs as, what that account can access, and where its credentials are stored.
  3. Check the high-exposure controls first. MFA on human accounts, network policies on the data platform, no secrets in repositories, and dependencies pinned to verified versions.
  4. Reduce privileges. Narrow every service account to the minimum it needs, and remove accounts that are no longer used.
  5. Turn on logging and review it. Make sure access to the data platform and the orchestration layer is logged, and that someone actually looks at unusual access patterns.

Infomineo transforms complex data into predictive insights and actionable intelligence through end-to-end analytics services, including data engineering and pipeline development and data quality assurance and validation.

Talk to our data analytics team →

Frequently Asked Questions

What makes a data pipeline secure?

A secure data pipeline controls who and what can access data at every stage: authenticated sources, encrypted transport, least-privilege service accounts, protected storage, governed sharing, and a build and deployment process with no secrets in code and pinned dependencies. Encryption alone is not enough, because most attacks arrive with valid credentials.

Is encryption enough to secure a data pipeline?

No. Encryption protects data from interception, but it does nothing against an attacker who logs in with stolen credentials or extracts secrets from a compromised build tool. The 2024 Snowflake customer breaches involved no encryption failure at all; the affected accounts lacked multi-factor authentication.

What is the difference between a secure data pipeline and a security data pipeline?

A secure data pipeline is any business data pipeline built with security controls. A security data pipeline is a specific category of tooling that collects and routes security logs and telemetry into monitoring platforms such as a SIEM. The terms sound alike but describe different things.

How should secrets be managed in a data pipeline?

Store every credential, API key, and connection string in a dedicated secrets manager that the pipeline reads at runtime, never in code or committed configuration. Rotate secrets on a schedule, scan repositories continuously for leaks, and treat any exposed secret as compromised.

How often should data pipeline access be reviewed?

Quarterly is a sensible default, and more often for pipelines handling regulated or sensitive data. Access accumulates as people change roles and projects end, and credentials often outlive the people and systems they were created for.

DATA ANALYTICS

Next-Gen Insights for Competitive Advantage

Infomineo transforms complex data into predictive insights and actionable intelligence through end-to-end analytics services, including data engineering and pipeline development and data quality assurance and validation. Top-tier strategy consulting firms and Fortune 500 companies partner with us for market intelligence, competitive analysis, and data-driven decision support, backed by 15 years of experience and 500,000+ client requests successfully completed.

Book A Discovery Call

WhatsApp