Running `resample` with `in_memory=False` reads all models into memory #8480

stscijgbot-jp · 2024-05-13T14:08:07Z

Issue JP-3621 was created on JIRA by Brett Graham:

ResampleStep when run with in_memory=False reads all models into memory. For a 33 member association the memory usage is shown in the attached graph.

One line that reads all models is:

jwst/jwst/resample/resample_step.py

Line 67 in 8b254ae

input = datamodels.open(input)

The text was updated successfully, but these errors were encountered:

stscijgbot-jp · 2024-08-02T14:33:11Z

Comment by Ned Molter on JIRA:

This should be fixed as part of implementing ModelLibrary (see JP-3690).

Here is a small summary of results from profiling memory for this step on its own in my PR branch, using as input an association containing a 46-cal-file subset of the data in

███████████████████████████████████████████

The total size of the input _cal files on disk is roughly 5 GiB.

Setting in_memory=True, the peak memory usage is 15.4 GiB, and I'm attaching a graph of usage over time to the ticket (resample_in_memory.png, sorry for the similar name to what Brett attached originally).

Setting in=memory=False, the peak memory usage is 10.6 GiB, and a graph is again attached (resample_on_disk.png).

This appears to be fixed. Analysis of the memory allocation breakdown from the flamegraph indicates that all the models are never loaded into memory. However, the most memory-intensive part of the step is resample_variance_arrays(), which requires making several large output arrays for the different types of variance that need to be resampled. In this example, each of those arrays has shape (9348, 13432). It is unclear to me how to improve this, and I'd say it's beyond the scope of this ticket.

stscijgbot-jp · 2024-09-20T14:33:27Z

Comment by Melanie Clarke on JIRA:

Fixed by #8683

stscijgbot-jp added the Software Affected: resample label May 13, 2024

emolter mentioned this issue Aug 12, 2024

JP-3690: Switch from ModelContainer to ModelLibrary for image3 pipeline #8683

Merged

8 tasks

tapastro closed this as completed in #8683 Sep 5, 2024

stscijgbot-jp added the performance-improvements label Oct 17, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Running `resample` with `in_memory=False` reads all models into memory #8480

Running `resample` with `in_memory=False` reads all models into memory #8480

stscijgbot-jp commented May 13, 2024

stscijgbot-jp commented Aug 2, 2024 •

edited

Loading

stscijgbot-jp commented Sep 20, 2024

Running resample with in_memory=False reads all models into memory #8480

Running resample with in_memory=False reads all models into memory #8480

Comments

stscijgbot-jp commented May 13, 2024

stscijgbot-jp commented Aug 2, 2024 • edited Loading

stscijgbot-jp commented Sep 20, 2024

Running `resample` with `in_memory=False` reads all models into memory #8480

Running `resample` with `in_memory=False` reads all models into memory #8480

stscijgbot-jp commented Aug 2, 2024 •

edited

Loading