Snowflake Data Loading: Bulk Load & Snowpipe | Workshop

Published: 05 April 2026
on channel: Surfalytics | Fast Track to Data Career
198
10

Surfalytics workshop: Snowflake data loading for analytics engineers — bulk load with `COPY INTO` and continuous ingestion with Snowpipe, with Azure-oriented demos and patterns that transfer to AWS and GCP.

📚 What you'll learn

Concepts
• Storage integration, external stage, and file format objects — one-time trust and reusable definitions
• `COPY INTO` anatomy: table, `@stage` path, patterns, file format, `ON ERROR`, `PURGE`, `VALIDATION_MODE`
• Load metadata and duplicate detection (~64 days bulk / 14 days Snowpipe context), `FORCE` reloads
• Bulk load orchestration: ad hoc runs, Snowflake tasks, Airflow, Azure Data Factory / APIs
• Snowpipe: event-driven ingestion (notifications, queues), serverless compute vs your warehouse
• Best practices: object-storage partitioning for bulk scans, file sizing (~100–250 MB guidance), parallelism vs warehouse threads, small-file overhead
• Cost & ops: LIST cost, notification charges, WH minimum billing (60s), serverless pricing multiplier, real-world tuning stories

Demos & discussion
• Live `COPY INTO`: validation mode, good vs bad files, partial loads, RESULTSCAN / error inspection, COPY_HISTORY / LOAD_HISTORY
• Snowpipe flow: notification integration, AUTO_INGEST, Azure Event Grid / queue wiring (conceptual + presupplied setup)
• Q&A: external tables + partition discovery costs, Kafka → S3 → Snowflake case, dbt external tables package
• Closing: storage partitioning vs Snowflake micro-partitions, cluster keys, Gamma slides + Excalidraw / Claude diagrams

⏱️ Timestamps
Full chapter list: see `youtube-timestamps.txt` in this folder.

📌 Resources mentioned
• Snowflake docs: COPY, Snowpipe, file formats, load history
• Select.dev / Definite (example blog on parallelism and file splitting)
• AWS Firehose–style buffering patterns (batch files before landing)
• Surfalytics companion: Terraform + Snowflake workshop (object management)

✅ Key takeaways
• Master `COPY INTO` first — Snowpipe reuses the same core load semantics with automation and different billing.
• Partition buckets and size files for predictable cost; many tiny files can dominate overhead and latency.
• Choose bulk (scheduled, your WH) vs Snowpipe (event-driven, serverless) using latency, ops burden, and $/credit math.
• Plan for deduplication windows and storage lifecycle after loads.
• Treat observability (history tables, tasks, alerts) as part of production ingestion design.

---
Surfalytics – Community workshop. Snowflake: Bulk Load & Snowpipe (Azure-focused demos).

#Surfalytics #Snowflake #DataEngineering #Snowpipe #COPYINTO #BulkLoad #Azure #BlobStorage #AnalyticsEngineering #ELT #DataLoading #CostOptimization #Workshop


On this page of the site you can watch the video online Snowflake Data Loading: Bulk Load & Snowpipe | Workshop with a duration of hours minute second in good quality, which was uploaded by the user Surfalytics | Fast Track to Data Career 05 April 2026, share the link with friends and acquaintances, this video has already been watched 198 times on youtube and it was liked by 10 viewers. Enjoy your viewing!