Profiling Lix with Tracy

10 Likes

Nice read, thank you! Great to see new work on profiling nix expressions around the ecosystem :slight_smile:

There’s been a PR for (cpp)nix using tracy as the profiler, before the sampling profiler ended up being merged in: Tracy-based expr evaluation profiler by picnoir · Pull Request #9967 · NixOS/nix · GitHub

The author, @picnoir wrote a comment about why he wouldn’t push this further after the sampling eval profiler in: Tracy-based expr evaluation profiler by picnoir · Pull Request #9967 · NixOS/nix · GitHub

Curious if you have seen that PR and have notes or thoughts on how it compares to your lix branch, especially regarding the complications about lazyness & forcing “the right” thunks & trace size on the other hand?

4 Likes

Nice!

Now that you mention it I remember reading about that at the time and didn’t use it much because I ran into the segfault issues mentioned. Since then I had completely forgotten about it though, so thanks for reminding me!

I’ve not seen this segfault issue at all, so maybe it’s been fixed in tracy or maybe it’s a difference in the functions used (I’m using ZoneScoped instead of ZoneTransient).

Regarding trace size, I think it really depends on how you are instrumenting things. You can definitely get unmanageably large traces if you are instrumenting everything. With my setup, the trace is about 4GB (according to the tracy UI) when evaluating the CI job at work. That’s been fine so far for me. While working on this, I had added traces to some functions and then removed them cause they led to too much data. I think it can be quite nice to just iterate yourself and see what works.

On lazy evaluation, that’s something I very briefly touched on in the post, but indeed it can be quite confusing. You need to have a clear understanding of lazy evaluation and how it intersects with your profiling tool. This hasn’t been a big issue for me because I didn’t try to get a very fine-grained understanding of evaluation from this tooling. It mostly pointed me to an area of the codebase and then my familiarity with the code made it quick to find the actual issue, eg, it looks like we are evaluating this bit of IFD twice, maybe it’s not being shared in a let-binding. I think my addition of the zone primop can help here as well cause you can scatter that into your code and see when stuff is being forced.

3 Likes

Hi there, amazing work! We do plan to add something similar to this (with some constraints on making it as much zero cost as possible but possible to enable dynamically) perhaps with Perfetto rather than Tracy (but those are quite similar) and offer a very rich ecosystem for profiling.

We also think we should have a OTEL layer for “higher” level spans, e.g. download speeds, IO stuff, etc.

6 Likes