Skip to content

tracing working group #671

Description

@Qard

@bnoordhuis @sam-github @othiym23 @wraithan @groundwater @brycebaril @trevnorris

We should revisit the tracing situation. Perhaps a working group is in order? Some of you have moved on from APM since our last discussion, some were not present, but your input is valuable.

Activity

  1. mikeal commented on Jan 30, 2015

    @mikeal
    Contributor

    quick question, is tracing part of the solution to a larger problem or is it totally separate?

    could tracing and dtrace support and systemtap support be classified as part of a larger effort the WG would be solving?

  2. mikeal commented on Jan 30, 2015

    @mikeal
    Contributor

    also @thlorenz who is working on some related stuff.

  3. Qard commented on Jan 30, 2015

    @Qard
    MemberAuthor

    I think they are all related. At the time of our previous discussion, AsyncListener was being removed from core. Now we have async_wrap, so I think we should reconvene to discuss what we can do with that to make tracing more pleasant. DTrace and SystemTap are just other destinations for trace data, so I think that could be part of the discussion.

    The topic of the working group could potentially be a bit broader, reaching into other debugging issues, like what to do with domains.

  4. trevnorris commented on Jan 30, 2015

    @trevnorris
    Contributor

    @Qard IMO the working group should discuss what tracing functionality is more generally wanted, and then the minimalist hooks to allow that functionality to live in user-land should be implemented here. There is a never ending list of features people want, and adding/maintaining all those is the wrong decision. Extending the hooks that currently live in async_wrap sounds great. I'd love to get more feedback on what more is needed.

  5. mikeal commented on Jan 30, 2015

    @mikeal
    Contributor

    I'm working on outreach right now to get wide input in to the roadmap. @trevnorris can you think of a good question to ask individuals and companies that could inform what they need from tracings? Maybe something like "what do you wish you knew about a running Node application that you don't know now?"

  6. Qard commented on Jan 30, 2015

    @Qard
    MemberAuthor

    In our previous meeting, we came to a rough consensus on what our problems were as APM providers. Knowing more about what the customers want to see would certainly be valuable too.

    I agree that a lot of the applications of this would likely be userland stuff, but we can figure out what that is when we get to it.

  7. trevnorris commented on Jan 30, 2015

    @trevnorris
    Contributor

    @mikeal Would it be possible for devs to provide pseudo examples of what they want? Meaning, they provide some example code and can explain what it is they want traced. Seeing usage cases, imho, would be the most beneficial.

  8. sam-github commented on Jan 30, 2015

    @sam-github
    Contributor

    @Qard do you have a link to the write up for that last meeting? I'm having trouble finding it, and its pretty relevant here.

  9. Qard commented on Jan 30, 2015

    @Qard
    MemberAuthor

    @sam-github Yep, here it is.

    https://gist.github.com/groundwater/942dad5c0c4cfae21af9

    These are the trace meeting notes compiled by @groundwater after our previous meeting, for anyone else that wants to have a look.

  10. brycebaril commented on Jan 30, 2015

    @brycebaril
    Contributor

    @Qard yes! Definitely still interested

  11. hayes commented on Jan 31, 2015

    @hayes

    I am definitely interested as well

  12. othiym23 commented on Jan 31, 2015

    @othiym23
    Contributor

    I am most definitely interested, both in following up on the conversation we started last year and in whatever broader working-group discussion that might result from Mikeal's canvassing.

  13. thlorenz commented on Jan 31, 2015

    @thlorenz
    Contributor

    I feel there are two goals here which are solved differently:

      1. improve tracing of the node process from within which in most cases requires either some sort of hooks, monkey patching, addons or changes to core
      1. simplify integration of system tool tracing which can be solved in user land in most cases

    I created an issue to collect info about existing user land tools, some of which are addons that interface with v8 in order to pull out profiling info.

    I'm mostly focusing on 2) ATM, which can be converters, parsers to consume output of perf, dtrace, system tap, etc. and plugins into tools like debuggers,

    At the moment we have the following (some rather new): I'm most likely missing lots, so please add some you know of

    For those interested, some of us are gathering in #ngin8 on IRC to discuss some of these efforts.
    @paulirish and I were talking there about _sunburst_s - (think flamegraphs+) - and how to get them integrated with system tools.

    So a wide scope to cover here, not sure if it makes sense to have this all in one group or if we should split it into one group for each section as outlined in 1) and 2).

  14. Qard commented on Jan 31, 2015

    @Qard
    MemberAuthor

    I see the working group mostly focusing on goal 1, but goal 2 ties into it in many ways.

    Being able to smoothly correlate trace data across JS and C++ boundaries is one thing that comes to mind. Visibility beyond the nebulous "it went into libuv somewhere, it'll probably call back at some point" would be great.

    Buffer usage in native modules also seems like something that'd be good to get some visibility into.

    I'm sure there's plenty of areas you all can think of where integrating with the native side could provide some very valuable data.

  15. 35 remaining items

  16. othiym23 commented on Feb 5, 2015

    @othiym23
    Contributor

    @piscisaureus I took notes and passed them on to @mikeal, who will submit them as a PR , and the meeting was recorded as a Hangout On Air, so it will be on the io.js channel.

  17. thlorenz commented on Feb 5, 2015

    @thlorenz
    Contributor

    I mentioned the problem about MachO symbols generated by v8.
    All that has gone into that is here

  18. mikeal commented on Feb 5, 2015

    @mikeal
    Contributor

    create a tracing-wg repo.

    meeting notes: https://github.com/iojs/tracing-wg/blob/master/wg-meetings/2015-02-05.md

    we should propose the next one there.

  19. natduca commented on Feb 6, 2015

    @natduca

    I thought I'd provide some notes on the direction we're taking in chrome and indirectly v8 around tracing & stack sampling. Since we both intersect at v8 and do run into some of the same "guhh i need a tracing api but its gotta be low overhead," I hope there's some useful context!

    All our performance tooling in chrome is moving to be trace-based. Just recently (chrome 42 maybe?) we moved devtools to use tracing for its timeline and flamechart-style js profiling. And, we've for a long time had chrome://tracing built around trace data. The rough philosophy we use for tracing is described in https://docs.google.com/a/chromium.org/document/d/1l0N1B4L4D94andL1BY39Rs_yXU8ktJBrKbt9EvOb-Oc/edit

    When we chromies say tracing, we mean a stream of events with timestamps. Beyond that, we've tried to embrace the notion that there is no One True Format To Rule Them All, but rather that the common clock ties everything together. As long as you can get a stream of traces from a few places, and have them share clocks or have clock alignment records, then the rest is a matter of postprocessing. Our UI (github.com/google/trace-viewer) for instance for chrome://tracing just merges kernel traces and userland traces after the fact.

    This is what allows us to have traces that contain OS/platform level data from dtrace/lttng/perf and also userland event streams. For instance, if you're on a chromebook, you can get a trace that has our userland trace data but also stuff from the kernel. Or, on android, similarly, using https://github.com/johnmccutchan/adb_trace. In both cases, we leave the OS-specific tracing to the OS, and then let the UI or tools then figure out how to pull our event stream and the OS' stream themselves.

    So, for most "day to day" tracing discussions, we solve all our stuff with c++ instrumentation. (Note, we consider sampling (e.g. interrupting every 1ms to get a stack trace) to be a userland activity.)

    Our userland tracing api is called "trace_event". This is a two layer thing, and actually resembles some of the ideas thrown around on the hangout I think: we have a header file with basic tracing macros, then this header can be configured to call into an implementation via a narrow set of tracing control apis. https://code.google.com/p/chromium/codesearch#chromium/src/base/trace_event/&q=trace_event&sq=package:chromium&type=cs ... though there are a few hundred trace apis that our code uses, it boils down to a half dozen c++ apis that one must implement in order to receive and process the events.

    This api is low enough overhead that we have thousands of trace events per second compiled into our shipping chrome builds, in our hottest chrome codepaths. We're pretty happy with the overhead: when its off, the overhead is just a check of a global char* for truthiness, as in, a few instructions [depending on architecture etc]. When its on, each tracepoint takes about 20-50us to buffer on a lowend android phone, which has been enough to keep us happy. On real machines with power, I'm pretty sure the overhead is way way lower. Shrug.

    In chrome, the actual trace api implementation is the trace_event_impl.cc, but in other libraries like Blink, Skia or V8, we copy the trace_event.h header file over and then apply a small tweak to it that causes it to trampoline back over to chrome via a local singleton. Then when we boot up one of those libraries in chrome, we set that singleton to point at chrome's tracing implementation. A key property of our architecture is that when a tracepoint is off, you don't have to hop from v8 to chrome in order to discover the trace isn't needed. I suspect this could let us trace things that are internal to v8 like promises without as much scary overhead discussions as we'd have if we were calling through an isolate-specific callback.

    This is all a bit wordy, I realize. It may be easier to read in code how we're hooking v8 up to chrome tracing... We are literally in the middle of connecting v8 to this tracing system. There are codereviews linked from that doc for context.

    https://docs.google.com/a/chromium.org/document/d/1_4LAnInOB8tM_DLjptWiszRwa4qwiSsDzMkO4tU-Qes/edit#

    Importantly, this isn't just being done for us to get v8 issue trace events for when its doing internal work, say compiling or doing a majorgc. We're also going to send sampling data stream plus code creation/relocation events through the trace interface, so that when we start tracing, we can get both timestampped "begin/end" events AND sampling events.

    One final technical note: I know one of the hot topics in all this is encoding the relationships between async information. In chrome's trace ecosystem, we've tried to solve this with flow events, which represent "arrows" between individual synchronous events on a thread, represented with regular begin/end events. More info on our trace data model --- which our trace_event.h file exists to generate --- is here: https://docs.google.com/document/d/1CvAClvFfyA5R-PhYUmn5OOQtYMH4h6I0nSsKchNAySU/edit

    Anyway, thats a lotta words. Hope it makes sense! This was all designed with client use cases heavily in mind, so I totally buy that this is all whacko and crazy when viewed through a server-side lens. This having been said, feel free to shout though with questions, I'm always around to chat.

  20. thlorenz commented on Feb 6, 2015

    @thlorenz
    Contributor

    Just wanted to point out that some of us interested in Profiling/Debugging are hanging out in IRC #ngin8.

  21. mikeal commented on Feb 6, 2015

    @mikeal
    Contributor

    @natduca we would love to have your input in the Tracing WG if you're interested :)

  22. sam-github commented on Feb 16, 2015

    @sam-github
    Contributor

    @natduca, re:

    in other libraries like Blink, Skia or V8, we copy the trace_event.h header file over and then apply a small tweak to it that causes it to trampoline back over to chrome via a local singleton.

    I can't find any sign of trace_event.h in https://chromium.googlesource.com/v8/v8 ... am I looking in the wrong place, or are you describing speculative future work?

  23. sam-github commented on Feb 16, 2015

    @sam-github
    Contributor
  24. natduca commented on Feb 17, 2015

    @natduca

    Yeah, this patch is stalled because the author has been busy. But its still coming. :)

  25. sam-github commented on Feb 17, 2015

    @sam-github
    Contributor

    @natduca can you comment on the differences and similarities between the Timeline view in Dev Tools, and the chrome://tracing view... do they both display trace_event data, or does only chrome:tracing show trace event data? Its a bit confusing, they seem very similar, perhaps one is destined to replace the other?

  26. thlorenz commented on Feb 17, 2015

    @thlorenz
    Contributor

    @sam-github Timeline in DevTools shows sampled data .cpuprofile, vs. chrome://tracing shows traced (structural) data.
    traceviewify actually emulates function entries/exits to convert .cpuprofiles to trace-viewer format.

  27. natduca commented on Feb 17, 2015

    @natduca

    @sam-github @thlorenz as in all big software, there's a lot of nuance because we're in between the old thing that is deprecated and the new thing that does't work quite right. :)

    Timeline view in devtools is almost completely based on chrome://tracing data these days. The corner cases are cpuprofile but we're working on moving that over completely. Same for network panel. The future we're working toward is that all performance data you see in devtools came through the tracing data stream.

    But, the UI are different: devtools ui is focused on ease of use for a web developer, whereas chrome://tracing is just the raw "good enough for chrome hackers" view. Devtools' timeline won't move to that chrome://tracing ui, though it may pick up ideas from that UI, or vice versa.

  28. paulirish commented on Feb 17, 2015

    @paulirish

    Pretty close, thorsten.

    differences and similarities between the Timeline view in Dev Tools, and the chrome://tracing view

    There's at least there UIs here:

    1. chrome://tracing aka trace-viewer - screenshot
    2. devtools timeline (flame chart) - screenshot
    3. devtools sampling profiler & flame chart - screenshot

    Once the above commit lands, all of these will be powered by the trace event data stream (v8 sampling and network being the remaining parts). And after that all 3 UIs will allow import/export of .trace files.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions