guess_batch_size#

anri.fwd.guess_batch_size(entries, hkls, F2, geom, row, det_shape, window=(3, 7, 7), mesh=None, memory_fraction=0.25)[source]#

Largest batch for render_row() whose render step fits in a fraction of the free memory.

Compiles the render step for two small batches (nothing is rendered) and reads XLA’s memory analysis to get the bytes per peak, then fits as many peaks as memory_fraction of the free memory allows: on GPUs, of the least free device; on CPU, of the host memory shared by the CPU devices, counting the host copy of each batch’s output too. Precision matters: float64 needs about twice the memory of float32. Selecting the peaks (select_chunk in render_row()) needs memory too, a few hundred MB by default, which is not counted.

Parameters:
Returns:

batch (int) – A power of two times the number of devices, at least 64 peaks per device