Repository navigation
proposal: a strategy to modularize core #7098
Description
Activity
- addeddiscussIssues opened for discussion and feedback.Issues opened for discussion and feedback.metaIssues and PRs related to the general management of the project.Issues and PRs related to the general management of the project.
on Jun 1, 2016 while I'm certainly in favor of the idea if a way can be found for it to work, I have deep concerns over the feasibility and usability. Another possible approach would be to have core modules published to npm in addition to being bundled, similarly to what we do with readable-streams. That, however, brings along it's own host of challenges and leads down a path leading straight to Java ClassLoader Hell. It's definitely worth exploring and discussing but I'm definitely skeptical.
@jasnell One big problem that we've seen with
readable-streamis that it'd be hard for node core to make changes to core modules that are then published to npm because many people will peg their dependency versions in ways (e.g. explicit versions or1.1.x) that would cause compatibility issues with node.One problem with splitting modules into separate repos would be keeping up with all of the PRs and issues and for users to be able to easily search across all of them. Imagine someone has a problem and so they go to search the issues/PRs of a particular module's repo where they think the problem lies and find nothing, despite there being an issue about the particular problem in a repo for a different core module.
- addedmoduleIssues and PRs related to the module subsystem.Issues and PRs related to the module subsystem.
on Jun 1, 2016 This basically requires to design whole new API layer, rewrite node from scratch and rebuild all modules on top of that. Doesn't seem realistic.
eljefedelrodeodeljefe commented
on Jun 1, 2016 ContributorMore actionsTo mitigate the native addons pain introducing a lightweight toolchain would be an option (there is a lot of precedence in languages). Relying on gyp and python and especially not having an an interface to them is the worst, currently. In my experience doing stuff with
clon NT is similarly pleasent as w/gcc. However this has only gotten easier the in the earlier last.
A strong point would be the coding style being forceably more self-contained. Currently I see a lot of intertwined (IMO) stuff that makes it hard for me, as a newcomer.My point is building shouldnt be the biggest obstacle.
This whole endeavour is as ambitious as it would be truly innovative.
@jasnell that is basically the approach I'm advocating. node would essentially become a curation of native add ons, when you install node in the usual way it would include a curated set of addons that are basically backwards compatible with what we have currently.
@mscdex you raise quite a good point. That is definitely a problem with github on very modularized projects. About the best they have is to search within an org.
Maybe we keep it in one repo on github? Would that have a significantly negative impact?
It might be a little tricky, but I think it's workable to have a workflow like now, where the latest is on master, and there are branches for the older versions of the parts.
Has anyone tried that on this scale before?
@Fishrock123 you could check in
node_modules, then you know exactly what you are getting.@eljefedelrodeodeljefe it has recently come to my attention that chrome (which gyp came from) has since moved on to ninja which looks like it might be that lightweight build too you are looking for.
Has anyone tried that on this scale before?
This is essentially what I did with the luvit 2.0 rewrite. I spent about a full year of my time working on the migration and rewrite and couldn't be happier with the result.
For those that don't know the history, luvit 1.0 was a re-implementation of node using Luajit and Lua instead of V8 and JS. It has the same APIs, a similar build system, and very similar architecture with a with luajit, libuv, openssl, etc compiled intto a single binary with the core modules embedded in the binary as compiled bytecode and available globally. It had a very node like command line interface and repl as well.
Working on the core modules was a nightmare and upgrading to newer versions of libuv was painful because of the deep coupling throughout the entire system. I had faithfully duplicated many of node's mistakes.
One problem with splitting modules into separate repos would be keeping up with all of the PRs and issues and for users to be able to easily search across all of them
In luvit, the luvit/luvit repo contains all the core modules that were part of luvit 1.0. They are in one repo, but published to lit (our equivalent to npm) as individual packages with a meta-package that depends on them all. They are versioned independently with proper interdependencies declared.
Having them all in one repo keeps things localized and this pattern has worked for other frameworks like the weblit framework (something like express or koa).
C addons are a massive pain in the butt as it is.
Yes they are. With the luvit 2.0 rewrite, I made two new projects at different levels of abstraction.
One is the luv project which is purely libuv bindings for lua. It can be used by other lua projects. It doesn't add any sugar (no emitters, no streams, no generators/coroutines, etc), but simply exposes libuv in as straightforward a manner as possible where it makes sense in the target language. (using lua multiple return values instead of C outargs for example).
I wish there was a V8 equivalent. That alone would be an awesome project and could live independent of node just like luv lives independent of luvit. For someone who knows both the V8 and libuv APIs, this is a smallish project. It took me a few weeks to write luv from scratch.
Then the next layer is known as luvi. It basically acts as the global C glue. In this single binary is contained a build of luajit, openssl, miniz (our zlib alternative), and luv. There is just enough glue code to bootstrap a module system that's implemented in userland in pure lua. Everything else is pure lua and not part of the C compile step at all.
Luvi has a sort of bundle virtual filesystem where it can work on a zip file or a directory and treat that tree as the core application. This abstraction means that you can work on core modules without any kind of build system whatsover. Also to make building the final binary as painless as possible, luvi looks for a zip file appended to itself at startup and uses that zip as the application.
So running the luvit application (what implements the old luvit 1.0 interface with core libs embedded) you can do it three ways.
# Run directly from disk passing in args luvi path/to/luvit -- args for luvit # Run directly from zip file luvi luvit.zip -- args for luvit # Build a binary cat `which luvi` luvit.zip > luvit chmod +x luvit ./luvit args for luvit
As you can see, once you have a luvi binary for your platform you don't need to ever touch a C compiler or linker again. Even if you're working on core modules, everything can be done without one.
I later added features in to the lit package manager / toolkit to do higher level things like automatically pulling in dependencies and building apps directly from urls (similar to go get).
There are still unsolved edges. The best way to include C modules it to build a custom luvi binary with your addons included, since the lua is modified more often than the C, you don't need a compiler for most work. The rackspace monitoring agent that uses luvit in production does this for some C code that does system diagnostics beyond what libuv can do.
I did add support for also bundling .so files in the zip or folder with clever tricks to dlopen them even in the case of zip file, but that still have the same portability problems that node currently faces with binary modules.
We did get one break since luajit has an excellent ffi and ctypes interface native to the VM. I don't know if we could bundle something similar with the v8/node equivalent to luvi, but it's saved me in many cases where I could do things like create ptys in pure lua or have efficient memory structures using ctypes.
Regarding http_parser, luvit 2.0 dumped it in favor of writing the parser in pure lua and depending on the JIT to make it fast. So far performance has not been a problem and we have one less C dependency to worry about.
V8 is pretty amazing at this kind of stuff as well and the cost of calling C++ from JS is much higher than the cost of calling C from Luajit so there is even more pressure to write more core in JS. To prove my point I once wrote a MD5 implementation in pure JS that was faster than node's openssl bindings to native MD5. The reason I won was two fold. I hit a sweet spot in the JIT and the overhead of calling C++ made openssl's version much slower than C++ alone.
Reacted by Venture, Josh Gavant, Tim Kevin Oxley and Jeremiah SenkpielReacted by Dominic Tarr and YoshReacted by Mikey, Josh Duff, Michiel Dral, Yosh and Rishi GiriReacted by Michael Jackson, Josh Duff, David Silverman, Sean Vieira and YoshSo the end result of all this work is one you install lit on your system (which is basically luvi + lit.zip concatenated), building the legacy luvit interface as a CLI app is simply:
lit make lit://luvit/luvit
It will sync down all the objects for the latest release of luvit (any objects seen before will be cached and already local), build a zip file, download (or load from cache) the appropriate luvi binary (possible extracting it's own if it matches the requirements) and building the luvit binary. No C compiler or linker needed!
This is exactly the same steps for building any other app. If you want a custom build of luvit with different builtin libraries, write your own, declare your dependencies and publish it. It will be installed the exact same way. If you want to make some other CLI app, use only the core deps you want or none at all. You can even swap out the require system for something custom if you want. The hooks in luvi are generic. I even found a clever way using shebang line to not include a copy of luvi into each binary, but a single line containing
#!/usr/local/bin/luvi --concatenated with the zip of bundled files.I wish node had this ability, how can we make it happen?
I'd like to nip this and ask you file an EP. It'll be much easier to follow if there's a single document that's kept up-to-date with the ongoing conversation.
12 remaining items
Exposing just C++ modules sounds like a nice goal, but I'm afraid it would be more difficult than it sounds. Since those C++ modules are meant for use with their JS counterparts, error checking is usually split: some things are easier to check in JS, and then in C++ you just assert, since they are not meant to be consumed directly.
Moreover, right now the C++ side could change in every possible way as long as it "looks" the same on JS land. Publishing current JS modules on npm implies making a promise on how the interaction with their C++ counterpart works, which leads to a broader API surface to maintain.
Couldn't be made the C++ part barenaked to the minimal functionality (mostly bindings to the library APIs) and move all the functionality to the Javascript modules? This probably would reduce the API surface...
@piranna Reduce?! It basically doubles it. Now you have two API surfaces, the JS side and the C++ side.
Touchée, I was thinking only about the C++ one.
@piranna I think your idea is a great way to eventually migrate people to a smaller core, but for the reasons @saghul mentioned, we first need to find a way to make the C bindings a supportable interface so that said modules published to npm can consume them.
I'm actively working on an experiment to build this from the ground up with the eventual goal to have enough userspace modules to re-implement node's core libraries. I'll expose things like libuv using their native API styles and make the core more friendly to bundling arbitrary JS modules so that nothing needs to be included by default.
I'm actively working on an experiment to build this from the ground up with the eventual goal to have enough userspace modules to re-implement node's core libraries. I'll expose things like libuv using their native API styles and make the core more friendly to bundling arbitrary JS modules so that nothing needs to be included by default.
Yeah, that's just what lead me here, I love the idea on your project! :-D
This issue has been inactive for sufficiently long that it seems like perhaps it should be closed. Feel free to re-open (or leave a comment requesting that it be re-opened) if you disagree. Just tidying up, not acting on a super-strong opinion or anything.
This issue has been inactive for sufficiently long that it seems like perhaps it should be closed. Feel free to re-open (or leave a comment requesting that it be re-opened) if you disagree. Just tidying up, not acting on a super-strong opinion or anything.
I'm still interested on this thing, but it's true that the discussion got stalled... Maybe the point is that nobody showed code here. Hope the issue gets open and the debate about how to do it start again.
eljefedelrodeodeljefe commented
on Feb 16, 2017 ContributorMore actions@mcollina might have done some efforts on the side.
- Reacted by Jesús Leganés-Combarro
eljefedelrodeodeljefe commented
on Feb 16, 2017 ContributorMore actionsIn my eyes @mcollina's efforts are or can be mature enough in the short term to actually have a practical implementation of it. So reopening would see appropriate to me, if Matteo wants to continue with it, since factually it seems to be the right thing to do/
I don't think discussing it here is the right place. Let's aim to land the streams EPS, and pull off that refactoring. I think it will work out well, and then we can move it forward by extracting other things.
Reacted by Jesús Leganés-CombarroI opened this more to start a discussion, which it did, node has gotta use it's issue tracker to actually manage things they are working on so okay to close this, will take this discussion to other venues
I don't actually think this is inactive in practice, however. Consider what we are doing for the
node debugsuccessor,node inspect: vendoring in thenode-inspectmodule.Truth is, anything like this is a pretty big step and not exactly a comfortable one to make. I think
node inspectwill probably help us learn a lot about vendoring a module in such a way, possibly the basis for doing more like it.
Developing modules within node has proven difficult, since user code does not specify which version of the core modules it depends on. This means that any breaking changes to those core modules may break an unknown number of node modules on npm and in people's applications.
The idea of removing most or all of the modules in npm was first floated (as far as I am aware) by @isaacs https://vimeo.com/56402326 and has since been kicked around. but how would we actually pull this off?
The first step would be just publishing the core modules on npm - many of the core modules are already published, often as browser shims, such as https://www.npmjs.com/package/buffer
but other times as basically unmaintained modules that like https://www.npmjs.com/package/net
I'm sure some thing could be figured out with the current owners of those modules.
Now, the tricky part is that many of the core modules depend on a part written in C.
C addons are a massive pain in the butt as it is. If you did
npm install httpand it compiled thehttp_parserthat would be a worse hell.so,
npm install httpshouldn't compile anything, and we don't want to break current node. how do we do this?What about, move the core modules into separate repos under the node org, before a release, npm install into node, and then node builds builds node/node_modules into the node binding statically.
This would also have the significant advantage that to distribute a node app that requires binary addons, you could build it into a custom statically linked node, and use that.
The modules should also be split into a -down part which contains the C part, and then an -up part which contains the javascript. The javascript usually changes a lot more than the C, we found this pattern in level and it works well.
When you require a core module, the same resolution process would apply as it does now, it will look in the node_modules folder, and then fall back to the list of core modules.
so, you could completely change the api in
httpbut if the C parthttp_parserwas backwards compatible, you wouldn't have to recompile anything. npm would just need to check what version of core modules the current instance of node has, and it would then not install those unless necessary.Or, those modules could just skip their build step if they detect they are already built into the running instance of node.
The part that I don't know about, would be setting up node's build script to remove change the hard coded deps to look in the
node/node_modulesinstead. C build system is not currently my strong suit, but surely there is a way this can be done?