Description
_ERROR_CODE_MAP (request_handler.py#L54-L68) maps 500 to ServerError and 503 to ServiceUnavailable. Both are in _SERVER_ERRORS, so the client marks the node unhealthy and retries on the next one. Every other 5xx, including 502 Bad Gateway and 504 Gateway Timeout, falls through to the generic TypesenseClientError (L318). That error goes straight to the caller: no retry, and the node stays marked healthy.
A load balancer or reverse proxy in front of Typesense answers 502 or 504 when its backend refuses the connection or does not answer in time. That is what happens while a node restarts. So with a proxy in front of the cluster or of each node, a restart turns into errors in the application, even though other nodes could have served the request.
typesense-js treats every status from 500 to 599 as ServerError and retries it. The 0.21.0 Python client had the same 500 and 503 only mapping, so this is longstanding rather than a 2.0.0 regression.
Steps to reproduce
No Typesense server is needed. For each status, a stub load balancer (nearest_node) answers that status and three stub nodes are healthy. The script makes one search per status.
pip install typesense==2.0.0
python3 repro_5xx.py
repro_5xx.py
"""
typesense-python 2.0.0: 502 and 504 responses are not retried and do not mark
the node unhealthy; only 500 and 503 are.
pip install typesense==2.0.0
python3 repro_5xx.py
For each status, the load balancer stub (nearest_node) answers that status
and nodes A, B, C are healthy. One search per status. A proxy or load balancer
in front of Typesense answers 502 or 504 when its backend refuses the
connection or times out, so these are the statuses a client sees while a node
restarts behind one.
"""
import asyncio
import collections
import importlib.metadata
import json
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import typesense
# --- Stub nodes: local HTTP servers that count each request on arrival, answer
# --- a fixed status and can be slow. No Typesense server is needed.
STATE = {} # port -> {"name", "latency", "status", "hits": Counter}
class _Stub(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def _serve(self):
st = STATE[self.server.server_address[1]]
length = int(self.headers.get("Content-Length") or 0)
if length:
self.rfile.read(length)
st["hits"][f"{self.command} {self.path.split('?')[0]}"] += 1
status = st["status"]() if callable(st["status"]) else st["status"]
if st["latency"]:
time.sleep(st["latency"])
if status == 200:
body = b'{"hits":[],"found":0}' if "search" in self.path else b"[]"
else:
body = json.dumps({"message": "Not Ready or Lagging"}).encode()
try:
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
except (BrokenPipeError, ConnectionResetError):
pass # the client gave up (timeout)
do_GET = _serve
do_POST = _serve
def log_message(self, *args):
pass
class _Server(ThreadingHTTPServer):
request_queue_size = 2048
daemon_threads = True
def stub_start(name, latency=0.0, status=200):
srv = _Server(("127.0.0.1", 0), _Stub)
port = srv.server_address[1]
STATE[port] = {"name": name, "latency": latency, "status": status, "hits": collections.Counter()}
threading.Thread(target=srv.serve_forever, daemon=True).start()
return port
def stub_node(port):
return {"host": "127.0.0.1", "port": str(port), "protocol": "http"}
def stub_hits(ports):
return {STATE[p]["name"]: sum(STATE[p]["hits"].values()) for p in ports}
def stub_health(client):
out = {"lb": client.config.nearest_node.healthy} if client.config.nearest_node else {}
for n in client.api_call.node_manager.nodes:
out[STATE[int(n.port)]["name"]] = n.healthy
return out
# --- end of stub nodes
SEARCH = {"q": "shoe", "query_by": "title"}
async def one(status):
ports = [stub_start("lb", 0, status), stub_start("A"), stub_start("B"), stub_start("C")]
client = typesense.AsyncClient({
"api_key": "xyz",
"nearest_node": stub_node(ports[0]),
"nodes": [stub_node(p) for p in ports[1:]],
})
try:
await client.collections["products"].documents.search(SEARCH)
outcome = "ok"
except Exception as e:
outcome = f"raised {type(e).__name__}"
await client.api_call.aclose()
h = stub_hits(ports)
retried = sum(v for k, v in h.items() if k != "lb") > 0
print(f"LB answers {status}: {outcome:<36} hits {h} retried on a node: {retried} LB healthy after: {client.config.nearest_node.healthy}")
return retried
async def main():
print("typesense", importlib.metadata.version("typesense"))
res = {s: await one(s) for s in (500, 502, 503, 504)}
reproduced = res[500] and res[503] and not res[502] and not res[504]
print("REPRODUCED: 500 and 503 fail over, 502 and 504 do not" if reproduced else "not reproduced")
return reproduced
if __name__ == "__main__":
raise SystemExit(0 if asyncio.run(main()) else 1)
Expected
All four statuses fail over to a healthy node and the search succeeds, as it does for 500 and 503.
Actual
typesense 2.0.0
LB answers 500: ok hits {'lb': 1, 'A': 1, 'B': 0, 'C': 0} retried on a node: True LB healthy after: False
LB answers 502: raised TypesenseClientError hits {'lb': 1, 'A': 0, 'B': 0, 'C': 0} retried on a node: False LB healthy after: True
LB answers 503: ok hits {'lb': 1, 'A': 1, 'B': 0, 'C': 0} retried on a node: True LB healthy after: False
LB answers 504: raised TypesenseClientError hits {'lb': 1, 'A': 0, 'B': 0, 'C': 0} retried on a node: False LB healthy after: True
REPRODUCED: 500 and 503 fail over, 502 and 504 do not
Suggested fix
Map the whole 5xx range to ServerError, keeping the specific 503 class:
@staticmethod
def _get_exception(http_code: int) -> typing.Type[TypesenseClientError]:
exception = _ERROR_CODE_MAP.get(str(http_code))
if exception is not None:
return exception
if 500 <= http_code <= 599:
return ServerError
return TypesenseClientError
Environment
- typesense-python 2.0.0 (tag
v2.0.0, e863b44). The same code is on master (ef3a3db).
- Python 3.12.14 (
python:3.12-slim).
Description
_ERROR_CODE_MAP(request_handler.py#L54-L68) maps 500 toServerErrorand 503 toServiceUnavailable. Both are in_SERVER_ERRORS, so the client marks the node unhealthy and retries on the next one. Every other 5xx, including 502 Bad Gateway and 504 Gateway Timeout, falls through to the genericTypesenseClientError(L318). That error goes straight to the caller: no retry, and the node stays marked healthy.A load balancer or reverse proxy in front of Typesense answers 502 or 504 when its backend refuses the connection or does not answer in time. That is what happens while a node restarts. So with a proxy in front of the cluster or of each node, a restart turns into errors in the application, even though other nodes could have served the request.
typesense-js treats every status from 500 to 599 as
ServerErrorand retries it. The 0.21.0 Python client had the same 500 and 503 only mapping, so this is longstanding rather than a 2.0.0 regression.Steps to reproduce
No Typesense server is needed. For each status, a stub load balancer (
nearest_node) answers that status and three stub nodes are healthy. The script makes one search per status.repro_5xx.pyExpected
All four statuses fail over to a healthy node and the search succeeds, as it does for 500 and 503.
Actual
Suggested fix
Map the whole 5xx range to
ServerError, keeping the specific 503 class:Environment
v2.0.0, e863b44). The same code is onmaster(ef3a3db).python:3.12-slim).