blog/gpu-pipeline #786
Replies: 1 comment
|
Hey, First of all, really amazing directions! This is majorly helpful. However, I have always had problems with chunking, especially along time, since with daily data, I typically have to interpolate. I understand this is very important, that is chunking, when the netcdf is global, but I am currently working with Brazilian ERA5. I found it better to batch-wise load and save the data as .zarrr. I have about 13000 daily netcdfs of Brazil, and I did this code. Stored it all as individual zarr files, and then work on that. The total time to save to zarr took me 3mins46sec import glob, xarray, os, tqdm, zarr all_paths = sorted(glob.glob("/home/darkmage/paper_three/vpd_netcdfs///*.nc")) batch_size = 1000 for i in tqdm.tqdm(range(0, total_files, batch_size)): |
Uh oh!
There was an error while loading. Please reload this page.
blog/gpu-pipeline
How to accelerate AI/ML workflows in Earth Sciences with GPU-native Xarray and Zarr.
https://xarray.dev/blog/gpu-pipeline
All reactions