Connecting Braze to Snowflake, Databricks and BigQuery with Cloud Data Ingestion

By the Stitch team · · 9 min read

Key takeaways

  • CDI supports five warehouses (Snowflake, Databricks, BigQuery, Redshift, Microsoft Fabric) plus S3, Azure Blob Storage and Google Cloud Storage.
  • Syncs are driven by an UPDATED_AT timestamp: get that column right and most problems go away.
  • Recurring syncs run every 15 minutes by default, down to every 5 minutes on request, and you can trigger syncs on demand.
  • CDI is a strong default when your warehouse is the source of truth. A reverse ETL tool earns its place when you need heavier transformation, orchestration or many destinations.
  • Design nested objects carefully. Every key in a nested object can count as a data point on legacy contracts, and payloads are capped.

Which warehouses does Braze Cloud Data Ingestion support?

According to Braze's CDI documentation, CDI connects to:

  • Cloud data warehouses: Amazon Redshift, Databricks, Google BigQuery, Microsoft Fabric and Snowflake
  • Cloud file storage: Amazon S3, Azure Blob Storage and Google Cloud Storage

Braze says CDI is available in all Braze regions and any Braze region can connect to any source data region, which matters if you're on AU-01 and your warehouse sits elsewhere.

Each warehouse has its own connection set-up:

  • Snowflake: key-pair authentication using a public key Braze provides, with a unique Snowflake user per Braze workspace. Auto-resume should be on for the warehouse.
  • Databricks: an OAuth service principal or a personal access token. Classic and Pro SQL warehouses can add two to five minutes of warm-up; serverless reduces this but "may result in slightly higher integration costs".
  • BigQuery: a service account, via Workload Identity Federation or a service account key, with access granted to the relevant datasets.

If your warehouse uses network policies, allowlist the Braze IPs for your instance.

How does a CDI sync work?

You create a table or view in your warehouse that Braze reads from. For the standard table set-up, each row needs:

  • UPDATED_AT: when the row was added or last changed
  • An identifier: EXTERNAL_ID, ALIAS_NAME with ALIAS_LABEL, BRAZE_ID, EMAIL or PHONE, with one identifier type per row
  • PAYLOAD: a JSON string of the fields to sync

On each run, Braze syncs rows where UPDATED_AT is later than the last value it synced. That makes syncs incremental, but only if your timestamps are honest.

Payload limits are worth knowing upfront. Braze says each row can carry a JSON object with up to 250 attributes, and payloads over 1 MB are rejected. By default, each run can sync up to 500 million rows.

What types of data can you sync?

CDI covers user and non-user data:

  • User attributes, including nested custom attributes, arrays of objects and subscription statuses
  • Custom events, with name required and time, app_id and properties optional
  • Purchase events, with product_id, currency and price required
  • Subscription states, as subscription_group_id and subscription_state pairs
  • Catalog items, for product, content or offer catalogs
  • User deletion requests

Field-level requirements for events, purchases and subscriptions are on the table set-up page. If an event row has no time, Braze uses UPDATED_AT as the event time. Braze's best practices say you can delete users by external ID, user alias or Braze ID.

Braze also offers zero-copy options. According to its comparison of ingestion options, CDI Segments and CDI Canvas triggers keep warehouse data in place rather than writing it to profiles. Those queries run in your warehouse, so you pay the compute, but Braze doesn't log data points for them.

How often can CDI sync?

Braze says recurring syncs can run as often as every 5 minutes or as rarely as once per month. The dashboard default minimum is 15 minutes, and 5-minute syncs need Braze Support or your customer success manager.

You can also trigger a sync through the API once a warehouse job finishes, which is often the cleanest pattern: your dbt or orchestration run completes, then Braze pulls the new rows. Only one sync can run per integration at a time.

Two other things to plan for:

  • Rate limits: CDI shares the rate limit with the Braze API, so heavy syncs can compete with real-time /users/track traffic.
  • Real-time needs: if something must hit Braze within seconds, such as a booking or a password reset, send it via the SDK or REST API. Braze describes /users/track as near real time.

How do you get Braze data back into the warehouse?

The return path matters as much as ingestion. You need engagement data in the warehouse to measure lift, build attribution and feed models.

  • Snowflake Data Sharing: Braze shares a read-only database straight into your Snowflake account. What you get depends on your Data Distribution entitlement: engagement events, user behaviour events, or Profile 360 with profile and attribute changes. Shared data uses no storage in your account, so you pay only for compute when you query it.
  • Currents: Braze's real-time stream of engagement events, delivered in Avro format. The available Currents partners for storage are Amazon S3, Google Cloud Storage and Azure Blob Storage, so for Databricks or BigQuery you'd typically land files in storage and load from there. Braze says a Currents connector is included in many pro and enterprise-level packages.
  • Newer options: In June 2026 Braze announced User Profile Streaming and a Delta Sharing integration (beta) for Databricks that sends engagement signals back to the warehouse.

When should you use CDI instead of a reverse ETL tool like Hightouch or Census?

Braze itself says many reverse ETL use cases can also be handled with native CDI, and that the choice depends on transformation, orchestration and data ownership.

Here's how we'd frame it:

Use CDI when:

  • Your warehouse is the source of truth and your data team can shape tables for Braze
  • Braze is the main or only destination
  • You want fewer tools and contracts to manage

Consider a reverse ETL tool like Hightouch or Census when:

  • You need to send the same audiences to Braze and to ad platforms, CRM or support tools
  • Marketers want a visual audience builder over warehouse models
  • You want change detection handled for you. Hightouch, for example, says it only syncs data that needs updating and sends only changed columns, though it notes this can be slower than overwriting

Plenty of stacks use both: CDI for core profile and event data, and a reverse ETL tool for cross-channel audiences.

What are the common CDI pitfalls?

What does CDI look like in a real implementation?

Stitch was lead Braze partner for Serko AI, working alongside Braze's onboarding team. Serko's data flows from its data warehouse into Braze via Cloud Data Ingestion, alongside the Braze Web SDK and a server-side integration via GTM.

Multi-leg trips were the tricky part. We co-designed a flattened active_flights array schema with Braze to work around nesting limits, then used Liquid to personalise messages with complex flight, hotel and cost objects. The lifecycle layer shipped to production in six weeks from kick-off, with booking confirmation, disruption SMS and 24-hour reminder flows active from day one.

How Stitch can help

Stitch is an independent Auckland consultancy and a Braze Alloys Solutions Partner (Orbit tier), named ANZ Rising Star of the Year at Braze's inaugural ANZ Partner Awards in 2026. We're also a Hightouch Implementation Partner and Segment Certified Partner, so we can help you pick the right ingestion pattern rather than defaulting to one tool.

We design warehouse schemas for Braze, set up CDI and the return path, and build the lifecycle programmes that use the data. Typical implementations take 8 to 12 weeks.

FAQ

Does Braze Cloud Data Ingestion support Databricks and BigQuery? Yes. Braze lists Snowflake, Databricks, Google BigQuery, Amazon Redshift and Microsoft Fabric as supported warehouses, plus S3, Azure Blob Storage and Google Cloud Storage.

How often can Braze CDI sync? Every 15 minutes by default, down to every 5 minutes if you ask Braze Support, or as rarely as monthly. You can also trigger a sync via the API after a warehouse job finishes.

What is the UPDATED_AT column for? Braze uses it to find rows added or changed since the last sync. If it's wrong, rows get skipped or re-synced, so keep it accurate and in UTC.

Can CDI delete users in Braze? Yes. Braze supports user deletion requests through CDI, using external ID, user alias or Braze ID.

Do I need Hightouch or Census if I have CDI? Not always. CDI covers many warehouse-to-Braze use cases. A reverse ETL tool helps when you need many destinations, visual audience building or managed change detection.

Sources

Talk to a Braze partner

More on Braze

Let's connect the dots.
Get in touch.

Or book a time that suits you. No pitch deck, no obligation.