Stuck on an issue?

Lightrun Answers was designed to reduce the constant googling that comes with debugging 3rd party libraries. It collects links to all the places you might be looking at while hunting down a tough bug.

And, if you’re still stuck at the end, we’re happy to hop on a call to see how we can help out.

Saving larger than memory output to HDF5

See original GitHub issue

Hi,

I am trying to process a dataset larger than memory in the following way. data is an HDF5 dataset.

>> def filter:
>>     ...
>>
>> daskdata= da.from_array(data,chunks=(300,400,1000))
>> output = daskdata.map_blocks(filter).compute(get=multiprocessing.get)

My problem now is output will be larger than memory. How can I avoid dumping output into memory? And the filter will return a numpy array.

Thanks

Issue Analytics

State:
Created 6 years ago
Comments:14 (8 by maintainers)

Top GitHub Comments

1reaction

jakirkhamcommented, Dec 5, 2017

Might try using h5pickle to workaround this.

Edit: Please be very careful when writing to HDF5 files. They are not designed to be written to from multiple processes. Would want to ensure they are written to from only one process at a time.

1reaction

mrocklincommented, Aug 11, 2017

Yes. Consider the use of the dask.set_options(get=...) context manager.

In the future, it would be good to see usage questions like this on stack overflow under the #dask tag