Skip to content

[Bug] HTTP 422 "invalid request error" on /provider/v1/chat/completions — rejects a 216,540-token prompt (below 256 Ki), is non-deterministic at ~212k, and caps minimal-field requests at prompt_tokens + max_completion_tokens ≥ 646,155 #952

Description

@UNscientific-9

Summary

Reproducible. Still failing at the time of writing: 2026-09-29 09:39 UTC. POST https://api.commandcode.ai/provider/v1/chat/completions with model deepseek/deepseek-v4.1-flash returns HTTP 422 invalid request error with no detail — no param, no code, no length information, only a trace_id — and the client request is a valid OpenAI chat-completions body.

Three measurements on 2026-09-29:

1. Our client's request shape is rejected far below 262,144 (256 Ki) prompt tokens. In session session-bb836559-… (contextWindow: 1000000, maxTokens: 384000) the endpoint returned 200 for a 215,824-token prompt at 09:38:06 UTC, then 422 for a ≈216,540-token prompt at 09:39:21 UTC — 75 seconds later, same field set, same max_completion_tokens: 384000. Three consecutive attempts (09:39:21 / 09:39:35 / 09:39:43 UTC) returned 422 and the session stopped responding. 216,540 prompt tokens is 82% of 262,144: the session never reached 256 Ki.

2. The same shape also failed at ≈212k, and an identical size succeeded 34 seconds later. 09:35:50 UTC → 422; 09:36:24 UTC → 200 with prompt_tokens 212,336. The two prompts are 13 tokens apart.

3. With a minimal field set (model, messages, max_completion_tokens only) the accepted/rejected boundary tracks prompt_tokens + max_completion_tokens. 200 up to 644,984; 422 from 646,155. A 380,012-token prompt with a reservation of 16 returned 200. Direct probes, same route, same key, 08:57–08:59 UTC today.

Impact. With max_completion_tokens: 384000 the largest prompt we ever got through was 260,984 tokens (probe) and 259,269 tokens (real session). Two real sessions died at that point and did not recover: session-784fcb42-… at 259,269 (08:38 UTC) and session-bb836559-… at 215,824 (09:38 UTC). Your documentation advertises up to 1M context for this route. Clients cannot recover, because the error body names no limit (see Questions 8).

Timeline — 2026-09-29, all times UTC, all commandcode + deepseek/deepseek-v4.1-flash, same API key

Time Session prompt_tokens max_completion_tokens HTTP trace_id
08:38:32 784fcb42 259,269 (your usage) 384,000 200 —
08:38:49 784fcb42 ≈262,155 384,000 422 f9965636e3da815215348ef4ddf2ae63
08:39:20 784fcb42 ≈262,155 384,000 422 8d2b78e48b922f3e14df82ca87c33073
08:57–08:59 probe P-B 260,984 384,000 200 —
08:57–08:59 probe EX-1 ≈263,050 384,000 422 b24db3978e79eba528520556aad60aac
08:57–08:59 probe P-A ≈263,050 16 422 691e97aac1847cba8ba8b3a7e011b03f
09:14:08 bb836559 211,271 384,000 200 —
09:35:50 bb836559 ≈212,323 384,000 422 de1f2c8d39f9ee0b5961a261772e06d4
09:36:24 bb836559 212,336 384,000 200 —
09:36:45 bb836559 214,915 384,000 200 —
09:37:52 bb836559 215,096 384,000 200 —
09:38:06 bb836559 215,824 384,000 200 —
09:39:21 bb836559 ≈216,540 384,000 422 789fac7f2c68d85e78c95cbe60a92bdc
09:39:35 bb836559 ≈216,540 384,000 422 5aab8b1aab4f8a81f2c8dbfd1dffcdc4
09:39:43 bb836559 ≈216,540 384,000 422 7d4ff01156fe25b243b1df2175d0d788

≈ = arithmetic from the preceding successful hop's provider-reported totals (rejected requests return no usage; see Measurement notes).

Probe results with a minimal field set:

prompt_tokens max_completion_tokens field set sum HTTP
380,012 16 minimal 380,028 200
249,980 384,000 minimal 633,980 200
260,984 (P-B) 384,000 minimal 644,984 200
≈263,050 (EX-1) 384,000 client field set ≈647,050 422
≈263,050 (P-A) 16 client field set ≈263,066 422

Environment

  • Endpoint: POST https://api.commandcode.ai/provider/v1/chat/completions (no proxy in the path)
  • Model id: deepseek/deepseek-v4.1-flash
  • Client: DeepSeek Harness 0.2.0-rc.1, route commandcode, api openai-completions, declared contextWindow 1,000,000 / maxTokens 384,000
  • Auth: standard Command Code API key (not disclosed)
  • The client's request shape: stream:true, stream_options:{include_usage:true}, store:false, reasoning_effort:"max", role:"developer" for the system turn, tools:[…]
  • All timestamps UTC

Steps to reproduce — we executed this and it reproduces

POST /provider/v1/chat/completions
Authorization: Bearer <key>
Content-Type: application/json

{
  "model": "deepseek/deepseek-v4.1-flash",
  "messages": [
    {"role": "developer", "content": "<~200 chars>"},
    {"role": "user", "content": "<filler: 1,126,980 chars of repeated text>"},
    {"role": "user", "content": "Reply with the single word OK."}
  ],
  "stream": true,
  "stream_options": {"include_usage": true},
  "store": false,
  "max_completion_tokens": 384000,
  "reasoning_effort": "max",
  "tools": [ /* any 1–2 tool schemas */ ]
}

→ HTTP 422, body ≈1.13 MB, ≈263,050 prompt tokens, content-type: application/json (not text/event-stream), no SSE events, 165-byte error body.

Contrast runs, all 200:

  1. Same filler, max_completion_tokens 384,000 → 16, no other field changes → 200
  2. Filler grown to 380,012 prompt tokens, reservation 16 → 200
  3. 249,980 prompt + 384,000 reservation, plain 3-message body (no stream/tools/developer) → 200
  4. Same as 3 but field named max_tokens instead → 200
  5. 260,984 prompt + 384,000 reservation, minimal fields → 200

Token counting on your side is deterministic, so prompt sizes are exact: prompt_tokens = 0.2333333 × filler_chars + 38. Calibrated on two points, then used to predict three further requests with 0-token error (249,980 / 270,014 / 380,012 all exact).

Observed

Real session session-784fcb42-…, 2026-09-29. 08:38:32 UTC success: your usage reports input 8,773 + cache read 250,496 = 259,269 prompt tokens, output 2,886 (total_tokens 262,155). 08:38:49 UTC next hop, prompt ≈ 262,155 → 422. 08:39:20 UTC the user typed "continue" → 422 again. Every later request in that session returned 422; it did not recover. Other sessions on the same key/provider kept working (97 successful calls after 08:38:49, max prompt 166,677).

Real session session-bb836559-…, 2026-09-29. Table above: 200 at 215,824 prompt tokens (09:38:06), 422 at ≈216,540 (09:39:21 / 09:39:35 / 09:39:43). The session stopped responding at 09:39:43 and did not recover. It had also been rejected once at ≈212,323 (09:35:50) and accepted 34 seconds later at 212,336 (09:36:24).

Raw 422 payloads — 8 responses, 8 distinct trace_ids:

422: {"message":"{\"message\":\"invalid request error trace_id: f9965636e3da815215348ef4ddf2ae63\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 8d2b78e48b922f3e14df82ca87c33073\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"error":{"message":"{\"message\":\"invalid request error trace_id: b24db3978e79eba528520556aad60aac\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}}
422: {"message":"{\"message\":\"invalid request error trace_id: 691e97aac1847cba8ba8b3a7e011b03f\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: de1f2c8d39f9ee0b5961a261772e06d4\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 789fac7f2c68d85e78c95cbe60a92bdc\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 5aab8b1aab4f8a81f2c8dbfd1dffcdc4\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 7d4ff01156fe25b243b1df2175d0d788\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}

No param, no code, no length information in any of them.

Client-side consequence. Our harness classifies a context-overflow error by matching message text (prompt is too long, context_length_exceeded, exceeds … maximum context length, etc.). None of these payloads contains any such phrase, so the error is classified as INVALID_REQUEST, not CONTEXT_WINDOW_EXCEEDED. The client therefore neither compacts nor retries: the session is permanently stalled at the point of failure.

Expected

HTTP 200, or an error that names the violated constraint — e.g. prompt_tokens + max_completion_tokens = 646155 > 645120 maximum. Your platform already returns the limit and the requested size elsewhere (cf. #911: 400 This model's maximum context length is 1048576 tokens. However, you requested …).

Evidence — control data, same provider + model id

Date (UTC) Successful calls with prompt > 262,144 Max prompt 422s
2026-09-27 60 371,246 none
2026-09-28 311 381,743 (last large success 352,480 @11:00:04) none
2026-09-29 0 259,269 (784fcb42 @08:38:32) 6 in real sessions + 2 probe 422s

The last observed successful call with prompt > 262,144 was at 2026-09-28 11:00:04 UTC. The first observed 422 was at 2026-09-29 08:38:49 UTC — ≈21.6 hours apart. On 2026-09-29 the largest successful prompt across all sessions was 259,269; session bb836559 was rejected at ≈216,540 with the same route and key.

What we ruled out

  • Prompt length alone — 380,012 prompt tokens returned 200, 45% past 262,144.
  • Output reservation alone — 249,980 + 384,000 = 633,980 returned 200.
  • The output field name — max_completion_tokens and max_tokens behave identically.
  • Request body byte size — 1.63 MB bodies returned 200; a 422 fired on a 1.13 MB body.
  • Account, quota, billing — 97 successful calls on the same key after the first failures, 5-hour window usage ≈3–4%, no 429 in history.
  • Client-side truncation or retry — we send the conversation as-is; each failing step was one POST, no retry. INVALID_REQUEST is not in our retry set.
  • Request content — no images; 65 tools/52 KB versus 66 tools/53 KB in a successful 09-28 request; flat prefix-cache hit rates across the boundary.
  • A different model id — all three days and all eight trace_ids are commandcode + deepseek/deepseek-v4.1-flash.
  • A fixed 262,144-token prompt ceiling — 380,012 prompt tokens returned 200; ≈216,540 returned 422.
  • A deterministic size threshold for our client's field set — 09:35:50 UTC ≈212,323 → 422; 09:36:24 UTC 212,336 → 200 (34 seconds later); 09:38:06 UTC 215,824 → 200; 09:39:21 UTC ≈216,540 → 422 (75 seconds later). Same key, model, client configuration, field set and max_completion_tokens.
  • One single prompt + max_completion_tokens ceiling for all request shapes — the minimal-field shape was accepted at 644,984 (08:57–08:59 UTC) and our client's shape was rejected at ≈600,540 (09:39:21 UTC) on the same day.

Measurement notes. Rejected requests return no usage, so the figures marked ≈ are arithmetic from the preceding successful hop's provider-reported totals (e.g. 215,824 prompt + 716 output → ≈216,540 for the next hop), not direct readings. Every other token figure in this report is a value reported by your API.

Questions for the maintainers

  1. What is the exact ceiling in force on deepseek/deepseek-v4.1-flash as of 2026-09-29 09:39 UTC, and what is it measured against — prompt tokens, prompt + max_completion_tokens, or body bytes?
  2. Is request validation on this route deterministic? Two requests 13 tokens apart (≈212,323 vs 212,336) returned 422 and 200 respectively, 34 seconds apart; two requests 716 tokens apart (215,824 vs ≈216,540) returned 200 and 422 respectively, 75 seconds apart. Same key, same model, same field set, same max_completion_tokens: 384000.
  3. Does the validator assert prompt_tokens + max_completion_tokens ≤ context_window? What is that window for this model — 645,120 (630 Ki)? 646,144 (631 Ki)? Our minimal-field probes bracket it to [644,984, 646,155).
  4. Did the limit for this route change on 2026-09-29? The same key was accepted with prompt + max_completion_tokens = 644,984 (minimal fields, 08:57–08:59 UTC) and rejected with ≈600,540 (our field set, 09:39:21 UTC).
  5. Is reasoning_effort: "max" accepted on this route? OpenAI's enum is minimal/low/medium/high/none; we send "max". Does the server use the max_completion_tokens value as sent, or a different output reservation, when reasoning_effort is set? Same questions for role:"developer", store:false and stream_options. Probe P-A (prompt + reservation sum only 263,066, our field set) was rejected.
  6. Trace lookups — which constraint did each of these violate?
    • f9965636e3da815215348ef4ddf2ae63 — 09:38:49 (prompt ≈262,155, reservation 384,000)
    • 8d2b78e48b922f3e14df82ca87c33073 — 09:39:20 (prompt ≈262,155, reservation 384,000)
    • b24db3978e79eba528520556aad60aac — 08:57–08:59 (probe EX-1)
    • 691e97aac1847cba8ba8b3a7e011b03f — 08:57–08:59 (probe P-A, reservation 16, sum 263,066)
    • de1f2c8d39f9ee0b5961a261772e06d4 — 09:35:50 (prompt ≈212,323, reservation 384,000)
    • 789fac7f2c68d85e78c95cbe60a92bdc — 09:39:21 (prompt ≈216,540, reservation 384,000)
    • 5aab8b1aab4f8a81f2c8dbfd1dffcdc4 — 09:39:35 (same)
    • 7d4ff01156fe25b243b1df2175d0d788 — 09:39:43 (same)
  7. Is max_completion_tokens: 384000 a valid value on this route? What value should clients send?
  8. Could this error name the violated limit and the observed size, and return 400 like your other oversized-prompt errors? With the current opaque body, clients cannot distinguish "reduce the reservation and continue" from "the request is malformed", and sessions stay stuck.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions