Skip to content

Fix suspend resolvers hanging when awaiting a data loader - #831

Open
oryan-block wants to merge 1 commit into
masterfrom
bugfix/419
Open

oryan-block wants to merge 1 commit into
masterfrom
bugfix/419

Conversation

@oryan-block

Copy link
Copy Markdown
Collaborator

Fixes #419

Checklist

  • Pull requests follows the contribution guide
  • New or modified functionality is covered by tests

Description

A suspend resolver that calls DataLoader#load or loadMany and then awaits the result hangs forever. This still happens on master (graphql-java 26.1) with default options, for root and nested fields. MethodFieldResolverDataFetcher#get started suspend resolvers with the default coroutine start, which hands the body to the configured dispatcher (Dispatchers.Default unless configured otherwise) and returns the future straight away. graphql-java dispatches a level's DataLoaders once that level's data fetchers have returned, so the load usually got queued after the dispatch had already run, and the await waited for a batch that never ran. It's basically the supplyAsync example the graphql-java batching docs tell you not to do, as someone pointed out on the issue.

The coroutine now starts with CoroutineStart.UNDISPATCHED, so the resolver runs on the calling thread until its first suspension. Loads are queued before the data fetcher returns and the rest of the coroutine still resumes on the configured context. An undispatched coroutine runs its body even if its context is already cancelled, which a default start never does, so the block calls ensureActive() first. That way a resolver still isn't invoked when, for example, the Job passed through SchemaParserOptions.coroutineContext is cancelled. Calling dispatch() by hand, like the workaround on the issue does, shouldn't be needed anymore for this case.

SuspendFunctionDataLoaderTest covers a root suspend field awaiting loadMany and a nested suspend field awaiting load. Each runs 20 executions with a 5s timeout so a hang fails the test instead of blocking the build. Both time out on master. There's also a new test in MethodFieldResolverDataFetcherTest for the cancelled context.

This only covers loads issued before the resolver first suspends. A load issued after a real suspension, like a second load that depends on the first or a load after withContext(Dispatchers.IO), still hangs with default options, on master and with this change. A plain fetcher that chains CompletableFuture loads hangs the same way, so I didn't try to work around it here. For those, graphql-java's DataLoaderDispatchingContextKeys.ENABLE_DATA_LOADER_CHAINING or ENABLE_DATA_LOADER_EXHAUSTED_DISPATCHING work together with this change. Chaining only helps loaders that come from env.getDataLoader, not ones held on a custom context object. Dispatchers.Unconfined isn't a full workaround either. It still hangs on master for suspend fields nested under another suspend field. Subscriptions are left alone.

Behaviour change: code in a suspend resolver up to its first suspension, including the resolver method call itself, now runs on the graphql-java calling thread instead of the dispatcher from SchemaParserOptions.coroutineContext / coroutineContextProvider. After the first suspension it resumes on the configured context as before, and the coroutine context itself is unchanged. So a suspend function that does blocking work without ever suspending (e.g. blocking JDBC with no withContext) now blocks the calling thread, same as a non-suspend resolver. Wrapping it in withContext(Dispatchers.IO) moves it off again. Exceptions thrown before the first suspension still fail the field the same way as before.

🤖 Generated with Claude Code

Suspend resolver methods were started with the default coroutine start,
which dispatches the body to the configured dispatcher and returns the
future straight away. graphql-java dispatches the DataLoaders of a
level once the data fetchers of that level have returned, so a resolver
that awaited DataLoader.load or loadMany often registered its keys only
after that dispatch had happened and then waited forever for a batch
that never ran.

Start the coroutine with CoroutineStart.UNDISPATCHED so the resolver
runs on the calling thread until its first suspension. Loads are then
queued before the data fetcher returns, as graphql-java expects, and
the rest of the coroutine still resumes on the configured context. An
undispatched coroutine runs even if its context is already cancelled,
so check for that first to keep cancelled resolvers from being invoked.

Loads issued after the resolver has suspended, such as a second load
that depends on the first, still rely on graphql-java's data loader
chaining or exhausted dispatching options, as with plain fetchers.

Fixes #419

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sonarqubecloud

sonarqubecloud Bot commented Oct 3, 2026

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Kotlin coroutine suspend fun resolver hangs when awaiting a data loader

1 participant