Cloud-optimized (CO) data formats are designed to efficiently store and access data directly from cloud storage without needing to download the entire dataset.
These formats enable faster data retrieval, scalability, and cost-effectiveness by allowing users to fetch only the necessary subsets of data.
They also allow for efficient parallel data processing using on-the-fly partitioning, which can considerably accelerate data management operations.
In this sense, cloud-optimized data is a nice fit for data-parallel jobs using serverless.
FaaS provides a data-driven scalable and cost-efficient experience, with practically no management burden.
Each serverless function will read and process a small portion of the cloud-optimized dataset, being read in parallel directly from object storage, significantly increasing the speedup.
In this talk, you will learn how to process cloud-optimized data formats in Python using the Lithops toolkit.
Lithops (https://github.com/lithops-cloud/lithops) is a serverless data processing toolkit that is specially designed to process data from Cloud Object Storage using Serverless functions.
We will also demonstrate the Dataplug library (https://github.com/CLOUDLAB-URV/dataplug) that enables Cloud Optimized data managament of scientific settings such as genomics, metabolomics, or geospatial data. We will show different data processing pipelines
in the Cloud that demonstrate the benefits of cloud-optimized data management.
On this page of the site you can watch the video online Giménez & Lopez - Processing Cloud optimized data in Python with Serverless Functions | SciPy 2025 with a duration of hours minute second in good quality, which was uploaded by the user SciPy 08 September 2025, share the link with friends and acquaintances, this video has already been watched 136 times on youtube and it was liked by 1 viewers. Enjoy your viewing!