Skip to content

build: PGO-enabled release builds #1409

Description

@bnoordhuis

This is something that has been at the bottom of my TODO list for some time now. I don't seem to get around to it so I thought I'd file this issue in the hope that someone else might. :-)

I'd like to investigate doing release builds with profile-guided optimizations enabled. Strawman Makefile change:

diff --git a/Makefile b/Makefile
index e93e817..2deffda 100644
--- a/Makefile
+++ b/Makefile
@@ -257,6 +257,13 @@ release-only:
                exit 1 ; \
        fi

+.PHONY: pgo
+pgo:   BUILDTYPE=Release
+pgo:
+       $(MAKE) CFLAGS+="-fprofile-generate" CXXFLAGS+="-fprofile-generate"
+       $(MAKE) bench-all
+       $(MAKE) CFLAGS+="-fprofile-use" CXXFLAGS+="-fprofile-use"
+
 pkg: $(PKG)

 $(PKG): release-only

The idea is to build an instrumented binary first, run the benchmarks to collect profile data that tell the compiler where and what to optimize, then build the final binary using the data from the previous step.

Open questions / unresolved issues:

  • The benchmark suite consists primarily of micro-benchmarks. It's not very representative of real-world applications. We would need something better for PGO.
  • It would make the release process a lot slower, particularly on the ARM buildbots. We could of course do non-PGO builds on slow machines.
  • I couldn't get the 32 bits build to link when PGO and LTO are enabled, at least not with a 32 bits toolchain: it runs out of memory. Could be resolved for the ia32 buildbots by using a 64 bits toolchain (I suspect that's already the case) but for ARM, the only option is to cross-compile and that sucks.
  • PGO may penalize the uncommon case, i.e., it may regress performance for use cases that are not covered by our benchmarks.
  • PGO's net effect may be zero.

Activity

  1. added
    buildIssues and PRs related to Node.js builds or CI infrastructure.
    benchmarkIssues and PRs related to Node.js benchmarks and benchmarking infrastructure.
    on Apr 13, 2015
  2. rvagg commented on Apr 14, 2015

    @rvagg
    Member

    I think the biggest problem here is your first point - we don't have a good set of reality-based benchmarks. I've been itching to try and get a WG spun up to focus on this but it hasn't clicked so far (btw I'm not a benchmark person, I'm just interested in seeing this happen). For now @iojs/build have been discussing some regular benchmarking system for builds but this would still require a better benchmark suite.

    PGO's net effect may be zero.

    Do you have any numbers to share on experiments, or even a gut-feel on this? I've not had experience with PGO compiles.

  3. jbergstroem commented on Apr 14, 2015

    @jbergstroem
    Member

    @rvagg another issue is that we (@iojs/build) don't really cater for build permutations - so we can't uphold quality assurance (or if it even works across our architecture). It's slightly off topic to this PR, so I'll stop here. I'd also like to see results of this since the PGO tests I've tried in other software hasn't really shown any strong benefits. It's hard to measure though (as mentioned above).

  4. bnoordhuis commented on Apr 14, 2015

    @bnoordhuis
    MemberAuthor

    Do you have any numbers to share on experiments, or even a gut-feel on this? I've not had experience with PGO compiles.

    It was a while ago and I didn't properly benchmark it. I ran an x64 PGO build through some http_simple_auto benchmarks and it was about 7-9% faster after but not consistently so. http_simple is notoriously fickle though.

    The binary (after stripping) was a few 100 kB smaller though, so that may have helped. In retrospect, I should have profiled the before and after binaries with perf record and check where the differences are. Something for the next guy or gal. :-)

    By the way, I turned on -Wl,--gc-sections in the non-PGO build to remove dead code. It's currently disabled because of buggy toolchains but I think we should be able to safely turn it on again. IIRC, the issue was with gcc 4.4 and/or 4.5 in combination with particular binutils versions.

  5. YurySolovyov commented on Apr 14, 2015

    @YurySolovyov

    Maybe you can pull some benchmarks from projects like express or ws, or other popular(npm top 5/10/20?) modules to get more "real" examples

  6. bnoordhuis commented on Jun 26, 2015

    @bnoordhuis
    MemberAuthor

    /cc @nodejs/benchmarking - perhaps you can incorporate this into your roadmap? I'll close this issue.

  7. HyperHCl commented on May 12, 2016

    @HyperHCl

    to get more "real" examples

    Running profiling against such cases should also be preferred over internal benchmarks since they are more closer to 'real-life usage'. Actually Firefox has some very complicated PGO cases to mimic daily use cases.

    (What's the status of this issue now?)

  8. bnoordhuis commented on May 12, 2016

    @bnoordhuis
    MemberAuthor

    (What's the status of this issue now?)

    I don't believe anyone is or has been working on it.

  9. added a commit that references this issue on Sep 27, 2018
  10. added a commit that references this issue on Oct 3, 2018
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkIssues and PRs related to Node.js benchmarks and benchmarking infrastructure.buildIssues and PRs related to Node.js builds or CI infrastructure.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions