Summary
Since 2026-09-29, deepseek/deepseek-v4.1-flash on the provider plan rejects a significant share of requests with HTTP 422 and a generic body that names no field and no limit. The failures are size-correlated but not deterministic: they appear only once the prompt exceeds roughly 256K tokens, and requests of the same or larger size succeed seconds apart on the same connection settings. The same model and parameters handled 550K-608K token prompts without errors from 2026-09-25 to 2026-09-28. Impact: about 23% of requests above ~256K tokens fail (13 of 56 on 2026-09-29), and because the error looks like an invalid-request failure the client cannot tell it apart from a context overflow, so the turn is aborted and work has to be resumed manually. Long agent sessions are not reliable on this model right now.
Expected Behavior
A chat/completions request that fits the model's advertised context window returns 200 and a completion, as it did until 2026-09-28 for prompts up to 608,097 tokens. If the prompt genuinely cannot be served, the error is distinguishable: a specific status and a machine-readable cause (for example 400 with code=context_length_exceeded and the limit in the message, or 413), so a client can compress the conversation and retry instead of treating the failure as a malformed request. Requests of the same size behave consistently, regardless of which backend in the pool serves them.
Actual Behavior
The request returns HTTP 422. The body is a JSON object whose "message" field is itself a JSON-encoded string:
outer: {"message": "", "type": "server_error"}
inner (decoded): {"message":"invalid request error trace_id: <trace_id>","type":"invalid_request_error"}
The body is identical apart from the trace_id in every occurrence, and carries no field name, no limit and no limit value, so it is not possible to tell whether this is request validation or a routing/backend failure.
Observed on 2026-09-29, counting distinct requests (deduplicated by response id / trace id):
| Prompt size |
Succeeded |
Failed (422) |
| < 262,144 tokens |
1,292 |
0 |
| > 262,144 tokens |
43 |
13 |
Every rejection occurred in a session whose preceding successful request had already reached at least 265,955 tokens; no rejection occurred below 262,144 tokens; the largest prompt that succeeded that day was 330,799 tokens, so it is not a hard cutoff either.
Failures and successes interleave at the same size in the same session:
- 13:47:03 rejected (preceding successful prompt ~286,800 tokens)
- 13:47:18 succeeded, with a LARGER prompt (287,711 tokens)
- 16:11:40 succeeded, 299,301-token prompt
- 16:11:44 rejected, same session, four seconds later
- 14:18:53 a summarization request was rejected
- 14:19:16 the equivalent request succeeded, 23 seconds later
The behaviour is consistent with part of the traffic above ~256K tokens being served by a backend with a smaller context limit than the rest of the pool.
Steps to reproduce the issue
- Use the provider plan endpoint https://api.commandcode.ai/provider/v1/ with model
deepseek/deepseek-v4.1-flash, streaming (SSE), OpenAI-compatible chat completions.
- Send requests with max_tokens=384000 and reasoning_effort=max, text only, no images.
- Let the conversation grow past ~256K prompt tokens (roughly 262,144) and keep sending single requests.
- Observe: requests of the same size alternate between 200 and HTTP 422
invalid request error. Below ~256K tokens the same parameters never fail. Repeating a rejected request is usually enough to get a 200 on the next attempt, which is why this is not reproducible on every attempt.
Note: max_tokens=384000 has been accepted continuously since at least 2026-09-25, including on requests that succeeded on 2026-09-29, so no parameter is statically invalid.
Command Code Version
n/a - API report, not the CLI (client details in Additional context)
Operating System
macOS
Terminal/IDE
DeepSeek Harness 0.2.0-rc.2 (web GUI), openai-completions adapter
Shell
zsh
Session file (optional)
No response
Fix prompt (optional)
Provider gateway, /provider/v1/ chat completions for OpenAI-compatible clients. When a request cannot be served - because the prompt exceeds the context limit of the backend it would be routed to, or because that backend is temporarily unusable - return a distinguishable error instead of a generic 422 invalid_request_error: 400 with type=invalid_request_error and code=context_length_exceeded plus the limit in the message when the prompt is too large, and a retryable 503/529 when the backend is temporarily unusable. Also make routing consistent: either only route a model to backends that satisfy its advertised context window, or make the failure retryable so clients can recover. Check: send the same >256K-token prompt N times and confirm every response is either a success or a classified error - never an unclassified 422 - and that repeated attempts are stable.
Additional context
Environment
- Client: DeepSeek Harness (dsh) 0.2.0-rc.2, adapter openai-completions, streaming SSE
- Endpoint: https://api.commandcode.ai/provider/v1/
- Model: deepseek/deepseek-v4.1-flash (exact id as listed by your endpoint)
- Request parameters: max_tokens=384000, reasoning_effort=max, ~35 tool definitions
- Content: text only - the failing requests contain no images and no audio
- Context window configured client-side: 1,000,000 tokens
- All timestamps CEST (UTC+2)
Largest successful prompt per day (same model, same parameters)
- 2026-09-25: 552,262 tokens
- 2026-09-26: 550,442 tokens
- 2026-09-27: 554,026 tokens
- 2026-09-28: 608,097 tokens
- 2026-09-29: 330,799 tokens, with 13 requests rejected
Change observed on 2026-09-29: the failures begin the same day that new models appeared in the catalogue served to this account (for example deepseek/deepseek-v4.1-flash-fast and claude-sonnet-5-5). No client-side change was made to this route: same model id and same request parameters as 2026-09-25 through 2026-09-28, when prompts up to 608,097 tokens completed without errors. The real maximum context window served for deepseek/deepseek-v4.1-flash now appears to be on the order of 256K rather than the 1,000,000 tokens assumed here.
Trace IDs of the rejected requests
- 2026-09-29 01:16:58 ebeccdf9f447b92f41220911f8716881
- 2026-09-29 13:32:45 210a01b5c40db4b23f4072bcfdc03d41
- 2026-09-29 13:33:00 310ec7fef980d185b790161426e6dffd
- 2026-09-29 13:35:10 50da5368dd9152f1b92f0290e6e96055
- 2026-09-29 13:36:27 2f5b7eb22529a97a9f1c293dd5863fe7
- 2026-09-29 13:37:07 8cbc03b6d9c337c36686a28b1fabbbe1
- 2026-09-29 13:47:03 4da0bd59c98189521b1b7346bd4f2a5a
- 2026-09-29 13:47:06 268d2f9e881def6bba70c7890c234e08
- 2026-09-29 13:47:22 664e01bd3c75f7b335866e9b205f2c3d
- 2026-09-29 13:50:36 7ec00d3a8308343f197448dc096bae51
- 2026-09-29 14:18:57 6af735c8e2c7eb963e921961b54e8ff7
- 2026-09-29 16:11:44 2826b95dd58064df6004d3ce82d95752
- 2026-09-29 17:44:31 d0a9d8460d70088afa95de0264ded495
Additional request-level details for any of the trace IDs above can be provided on request.
Summary
Since 2026-09-29, deepseek/deepseek-v4.1-flash on the provider plan rejects a significant share of requests with HTTP 422 and a generic body that names no field and no limit. The failures are size-correlated but not deterministic: they appear only once the prompt exceeds roughly 256K tokens, and requests of the same or larger size succeed seconds apart on the same connection settings. The same model and parameters handled 550K-608K token prompts without errors from 2026-09-25 to 2026-09-28. Impact: about 23% of requests above ~256K tokens fail (13 of 56 on 2026-09-29), and because the error looks like an invalid-request failure the client cannot tell it apart from a context overflow, so the turn is aborted and work has to be resumed manually. Long agent sessions are not reliable on this model right now.
Expected Behavior
A chat/completions request that fits the model's advertised context window returns 200 and a completion, as it did until 2026-09-28 for prompts up to 608,097 tokens. If the prompt genuinely cannot be served, the error is distinguishable: a specific status and a machine-readable cause (for example 400 with code=context_length_exceeded and the limit in the message, or 413), so a client can compress the conversation and retry instead of treating the failure as a malformed request. Requests of the same size behave consistently, regardless of which backend in the pool serves them.
Actual Behavior
The request returns HTTP 422. The body is a JSON object whose "message" field is itself a JSON-encoded string:
outer: {"message": "", "type": "server_error"}
inner (decoded): {"message":"invalid request error trace_id: <trace_id>","type":"invalid_request_error"}
The body is identical apart from the trace_id in every occurrence, and carries no field name, no limit and no limit value, so it is not possible to tell whether this is request validation or a routing/backend failure.
Observed on 2026-09-29, counting distinct requests (deduplicated by response id / trace id):
Every rejection occurred in a session whose preceding successful request had already reached at least 265,955 tokens; no rejection occurred below 262,144 tokens; the largest prompt that succeeded that day was 330,799 tokens, so it is not a hard cutoff either.
Failures and successes interleave at the same size in the same session:
The behaviour is consistent with part of the traffic above ~256K tokens being served by a backend with a smaller context limit than the rest of the pool.
Steps to reproduce the issue
deepseek/deepseek-v4.1-flash, streaming (SSE), OpenAI-compatible chat completions.invalid request error. Below ~256K tokens the same parameters never fail. Repeating a rejected request is usually enough to get a 200 on the next attempt, which is why this is not reproducible on every attempt.Note: max_tokens=384000 has been accepted continuously since at least 2026-09-25, including on requests that succeeded on 2026-09-29, so no parameter is statically invalid.
Command Code Version
n/a - API report, not the CLI (client details in Additional context)
Operating System
macOS
Terminal/IDE
DeepSeek Harness 0.2.0-rc.2 (web GUI), openai-completions adapter
Shell
zsh
Session file (optional)
No response
Fix prompt (optional)
Provider gateway, /provider/v1/ chat completions for OpenAI-compatible clients. When a request cannot be served - because the prompt exceeds the context limit of the backend it would be routed to, or because that backend is temporarily unusable - return a distinguishable error instead of a generic 422 invalid_request_error: 400 with type=invalid_request_error and code=context_length_exceeded plus the limit in the message when the prompt is too large, and a retryable 503/529 when the backend is temporarily unusable. Also make routing consistent: either only route a model to backends that satisfy its advertised context window, or make the failure retryable so clients can recover. Check: send the same >256K-token prompt N times and confirm every response is either a success or a classified error - never an unclassified 422 - and that repeated attempts are stable.
Additional context
Environment
Largest successful prompt per day (same model, same parameters)
Change observed on 2026-09-29: the failures begin the same day that new models appeared in the catalogue served to this account (for example deepseek/deepseek-v4.1-flash-fast and claude-sonnet-5-5). No client-side change was made to this route: same model id and same request parameters as 2026-09-25 through 2026-09-28, when prompts up to 608,097 tokens completed without errors. The real maximum context window served for deepseek/deepseek-v4.1-flash now appears to be on the order of 256K rather than the 1,000,000 tokens assumed here.
Trace IDs of the rejected requests
Additional request-level details for any of the trace IDs above can be provided on request.