Skip to content

502 and 504 responses are not retried on another node (only 500 and 503 are) #145

Description

@alangmartini

Description

_ERROR_CODE_MAP (request_handler.py#L54-L68) maps 500 to ServerError and 503 to ServiceUnavailable. Both are in _SERVER_ERRORS, so the client marks the node unhealthy and retries on the next one. Every other 5xx, including 502 Bad Gateway and 504 Gateway Timeout, falls through to the generic TypesenseClientError (L318). That error goes straight to the caller: no retry, and the node stays marked healthy.

A load balancer or reverse proxy in front of Typesense answers 502 or 504 when its backend refuses the connection or does not answer in time. That is what happens while a node restarts. So with a proxy in front of the cluster or of each node, a restart turns into errors in the application, even though other nodes could have served the request.

typesense-js treats every status from 500 to 599 as ServerError and retries it. The 0.21.0 Python client had the same 500 and 503 only mapping, so this is longstanding rather than a 2.0.0 regression.

Steps to reproduce

No Typesense server is needed. For each status, a stub load balancer (nearest_node) answers that status and three stub nodes are healthy. The script makes one search per status.

pip install typesense==2.0.0
python3 repro_5xx.py
repro_5xx.py
"""
typesense-python 2.0.0: 502 and 504 responses are not retried and do not mark
the node unhealthy; only 500 and 503 are.

  pip install typesense==2.0.0
  python3 repro_5xx.py

For each status, the load balancer stub (nearest_node) answers that status
and nodes A, B, C are healthy. One search per status. A proxy or load balancer
in front of Typesense answers 502 or 504 when its backend refuses the
connection or times out, so these are the statuses a client sees while a node
restarts behind one.
"""

import asyncio
import collections
import importlib.metadata
import json
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

import typesense

# --- Stub nodes: local HTTP servers that count each request on arrival, answer
# --- a fixed status and can be slow. No Typesense server is needed.
STATE = {}  # port -> {"name", "latency", "status", "hits": Counter}


class _Stub(BaseHTTPRequestHandler):
    protocol_version = "HTTP/1.1"

    def _serve(self):
        st = STATE[self.server.server_address[1]]
        length = int(self.headers.get("Content-Length") or 0)
        if length:
            self.rfile.read(length)
        st["hits"][f"{self.command} {self.path.split('?')[0]}"] += 1
        status = st["status"]() if callable(st["status"]) else st["status"]
        if st["latency"]:
            time.sleep(st["latency"])
        if status == 200:
            body = b'{"hits":[],"found":0}' if "search" in self.path else b"[]"
        else:
            body = json.dumps({"message": "Not Ready or Lagging"}).encode()
        try:
            self.send_response(status)
            self.send_header("Content-Type", "application/json")
            self.send_header("Content-Length", str(len(body)))
            self.end_headers()
            self.wfile.write(body)
        except (BrokenPipeError, ConnectionResetError):
            pass  # the client gave up (timeout)

    do_GET = _serve
    do_POST = _serve

    def log_message(self, *args):
        pass


class _Server(ThreadingHTTPServer):
    request_queue_size = 2048
    daemon_threads = True


def stub_start(name, latency=0.0, status=200):
    srv = _Server(("127.0.0.1", 0), _Stub)
    port = srv.server_address[1]
    STATE[port] = {"name": name, "latency": latency, "status": status, "hits": collections.Counter()}
    threading.Thread(target=srv.serve_forever, daemon=True).start()
    return port


def stub_node(port):
    return {"host": "127.0.0.1", "port": str(port), "protocol": "http"}


def stub_hits(ports):
    return {STATE[p]["name"]: sum(STATE[p]["hits"].values()) for p in ports}


def stub_health(client):
    out = {"lb": client.config.nearest_node.healthy} if client.config.nearest_node else {}
    for n in client.api_call.node_manager.nodes:
        out[STATE[int(n.port)]["name"]] = n.healthy
    return out
# --- end of stub nodes

SEARCH = {"q": "shoe", "query_by": "title"}


async def one(status):
    ports = [stub_start("lb", 0, status), stub_start("A"), stub_start("B"), stub_start("C")]
    client = typesense.AsyncClient({
        "api_key": "xyz",
        "nearest_node": stub_node(ports[0]),
        "nodes": [stub_node(p) for p in ports[1:]],
    })
    try:
        await client.collections["products"].documents.search(SEARCH)
        outcome = "ok"
    except Exception as e:
        outcome = f"raised {type(e).__name__}"
    await client.api_call.aclose()
    h = stub_hits(ports)
    retried = sum(v for k, v in h.items() if k != "lb") > 0
    print(f"LB answers {status}: {outcome:<36} hits {h}  retried on a node: {retried}  LB healthy after: {client.config.nearest_node.healthy}")
    return retried


async def main():
    print("typesense", importlib.metadata.version("typesense"))
    res = {s: await one(s) for s in (500, 502, 503, 504)}
    reproduced = res[500] and res[503] and not res[502] and not res[504]
    print("REPRODUCED: 500 and 503 fail over, 502 and 504 do not" if reproduced else "not reproduced")
    return reproduced


if __name__ == "__main__":
    raise SystemExit(0 if asyncio.run(main()) else 1)

Expected

All four statuses fail over to a healthy node and the search succeeds, as it does for 500 and 503.

Actual

typesense 2.0.0
LB answers 500: ok                                   hits {'lb': 1, 'A': 1, 'B': 0, 'C': 0}  retried on a node: True  LB healthy after: False
LB answers 502: raised TypesenseClientError          hits {'lb': 1, 'A': 0, 'B': 0, 'C': 0}  retried on a node: False  LB healthy after: True
LB answers 503: ok                                   hits {'lb': 1, 'A': 1, 'B': 0, 'C': 0}  retried on a node: True  LB healthy after: False
LB answers 504: raised TypesenseClientError          hits {'lb': 1, 'A': 0, 'B': 0, 'C': 0}  retried on a node: False  LB healthy after: True
REPRODUCED: 500 and 503 fail over, 502 and 504 do not

Suggested fix

Map the whole 5xx range to ServerError, keeping the specific 503 class:

@staticmethod
def _get_exception(http_code: int) -> typing.Type[TypesenseClientError]:
    exception = _ERROR_CODE_MAP.get(str(http_code))
    if exception is not None:
        return exception
    if 500 <= http_code <= 599:
        return ServerError
    return TypesenseClientError

Environment

  • typesense-python 2.0.0 (tag v2.0.0, e863b44). The same code is on master (ef3a3db).
  • Python 3.12.14 (python:3.12-slim).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions