Summary
Reproducible. Still failing at the time of writing: 2026-09-29 09:39 UTC. POST https://api.commandcode.ai/provider/v1/chat/completions with model deepseek/deepseek-v4.1-flash returns HTTP 422 invalid request error with no detail — no param, no code, no length information, only a trace_id — and the client request is a valid OpenAI chat-completions body.
Three measurements on 2026-09-29:
1. Our client's request shape is rejected far below 262,144 (256 Ki) prompt tokens. In session session-bb836559-… (contextWindow: 1000000, maxTokens: 384000) the endpoint returned 200 for a 215,824-token prompt at 09:38:06 UTC, then 422 for a ≈216,540-token prompt at 09:39:21 UTC — 75 seconds later, same field set, same max_completion_tokens: 384000. Three consecutive attempts (09:39:21 / 09:39:35 / 09:39:43 UTC) returned 422 and the session stopped responding. 216,540 prompt tokens is 82% of 262,144: the session never reached 256 Ki.
2. The same shape also failed at ≈212k, and an identical size succeeded 34 seconds later. 09:35:50 UTC → 422; 09:36:24 UTC → 200 with prompt_tokens 212,336. The two prompts are 13 tokens apart.
3. With a minimal field set (model, messages, max_completion_tokens only) the accepted/rejected boundary tracks prompt_tokens + max_completion_tokens. 200 up to 644,984; 422 from 646,155. A 380,012-token prompt with a reservation of 16 returned 200. Direct probes, same route, same key, 08:57–08:59 UTC today.
Impact. With max_completion_tokens: 384000 the largest prompt we ever got through was 260,984 tokens (probe) and 259,269 tokens (real session). Two real sessions died at that point and did not recover: session-784fcb42-… at 259,269 (08:38 UTC) and session-bb836559-… at 215,824 (09:38 UTC). Your documentation advertises up to 1M context for this route. Clients cannot recover, because the error body names no limit (see Questions 8).
Timeline — 2026-09-29, all times UTC, all commandcode + deepseek/deepseek-v4.1-flash, same API key
| Time |
Session |
prompt_tokens |
max_completion_tokens |
HTTP |
trace_id |
| 08:38:32 |
784fcb42 |
259,269 (your usage) |
384,000 |
200 |
— |
| 08:38:49 |
784fcb42 |
≈262,155 |
384,000 |
422 |
f9965636e3da815215348ef4ddf2ae63 |
| 08:39:20 |
784fcb42 |
≈262,155 |
384,000 |
422 |
8d2b78e48b922f3e14df82ca87c33073 |
| 08:57–08:59 |
probe P-B |
260,984 |
384,000 |
200 |
— |
| 08:57–08:59 |
probe EX-1 |
≈263,050 |
384,000 |
422 |
b24db3978e79eba528520556aad60aac |
| 08:57–08:59 |
probe P-A |
≈263,050 |
16 |
422 |
691e97aac1847cba8ba8b3a7e011b03f |
| 09:14:08 |
bb836559 |
211,271 |
384,000 |
200 |
— |
| 09:35:50 |
bb836559 |
≈212,323 |
384,000 |
422 |
de1f2c8d39f9ee0b5961a261772e06d4 |
| 09:36:24 |
bb836559 |
212,336 |
384,000 |
200 |
— |
| 09:36:45 |
bb836559 |
214,915 |
384,000 |
200 |
— |
| 09:37:52 |
bb836559 |
215,096 |
384,000 |
200 |
— |
| 09:38:06 |
bb836559 |
215,824 |
384,000 |
200 |
— |
| 09:39:21 |
bb836559 |
≈216,540 |
384,000 |
422 |
789fac7f2c68d85e78c95cbe60a92bdc |
| 09:39:35 |
bb836559 |
≈216,540 |
384,000 |
422 |
5aab8b1aab4f8a81f2c8dbfd1dffcdc4 |
| 09:39:43 |
bb836559 |
≈216,540 |
384,000 |
422 |
7d4ff01156fe25b243b1df2175d0d788 |
≈ = arithmetic from the preceding successful hop's provider-reported totals (rejected requests return no usage; see Measurement notes).
Probe results with a minimal field set:
| prompt_tokens |
max_completion_tokens |
field set |
sum |
HTTP |
| 380,012 |
16 |
minimal |
380,028 |
200 |
| 249,980 |
384,000 |
minimal |
633,980 |
200 |
| 260,984 (P-B) |
384,000 |
minimal |
644,984 |
200 |
| ≈263,050 (EX-1) |
384,000 |
client field set |
≈647,050 |
422 |
| ≈263,050 (P-A) |
16 |
client field set |
≈263,066 |
422 |
Environment
- Endpoint:
POST https://api.commandcode.ai/provider/v1/chat/completions (no proxy in the path)
- Model id:
deepseek/deepseek-v4.1-flash
- Client: DeepSeek Harness 0.2.0-rc.1, route
commandcode, api openai-completions, declared contextWindow 1,000,000 / maxTokens 384,000
- Auth: standard Command Code API key (not disclosed)
- The client's request shape:
stream:true, stream_options:{include_usage:true}, store:false, reasoning_effort:"max", role:"developer" for the system turn, tools:[…]
- All timestamps UTC
Steps to reproduce — we executed this and it reproduces
POST /provider/v1/chat/completions
Authorization: Bearer <key>
Content-Type: application/json
{
"model": "deepseek/deepseek-v4.1-flash",
"messages": [
{"role": "developer", "content": "<~200 chars>"},
{"role": "user", "content": "<filler: 1,126,980 chars of repeated text>"},
{"role": "user", "content": "Reply with the single word OK."}
],
"stream": true,
"stream_options": {"include_usage": true},
"store": false,
"max_completion_tokens": 384000,
"reasoning_effort": "max",
"tools": [ /* any 1–2 tool schemas */ ]
}
→ HTTP 422, body ≈1.13 MB, ≈263,050 prompt tokens, content-type: application/json (not text/event-stream), no SSE events, 165-byte error body.
Contrast runs, all 200:
- Same filler,
max_completion_tokens 384,000 → 16, no other field changes → 200
- Filler grown to 380,012 prompt tokens, reservation 16 → 200
- 249,980 prompt + 384,000 reservation, plain 3-message body (no stream/tools/developer) → 200
- Same as 3 but field named
max_tokens instead → 200
- 260,984 prompt + 384,000 reservation, minimal fields → 200
Token counting on your side is deterministic, so prompt sizes are exact: prompt_tokens = 0.2333333 × filler_chars + 38. Calibrated on two points, then used to predict three further requests with 0-token error (249,980 / 270,014 / 380,012 all exact).
Observed
Real session session-784fcb42-…, 2026-09-29. 08:38:32 UTC success: your usage reports input 8,773 + cache read 250,496 = 259,269 prompt tokens, output 2,886 (total_tokens 262,155). 08:38:49 UTC next hop, prompt ≈ 262,155 → 422. 08:39:20 UTC the user typed "continue" → 422 again. Every later request in that session returned 422; it did not recover. Other sessions on the same key/provider kept working (97 successful calls after 08:38:49, max prompt 166,677).
Real session session-bb836559-…, 2026-09-29. Table above: 200 at 215,824 prompt tokens (09:38:06), 422 at ≈216,540 (09:39:21 / 09:39:35 / 09:39:43). The session stopped responding at 09:39:43 and did not recover. It had also been rejected once at ≈212,323 (09:35:50) and accepted 34 seconds later at 212,336 (09:36:24).
Raw 422 payloads — 8 responses, 8 distinct trace_ids:
422: {"message":"{\"message\":\"invalid request error trace_id: f9965636e3da815215348ef4ddf2ae63\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 8d2b78e48b922f3e14df82ca87c33073\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"error":{"message":"{\"message\":\"invalid request error trace_id: b24db3978e79eba528520556aad60aac\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}}
422: {"message":"{\"message\":\"invalid request error trace_id: 691e97aac1847cba8ba8b3a7e011b03f\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: de1f2c8d39f9ee0b5961a261772e06d4\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 789fac7f2c68d85e78c95cbe60a92bdc\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 5aab8b1aab4f8a81f2c8dbfd1dffcdc4\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
422: {"message":"{\"message\":\"invalid request error trace_id: 7d4ff01156fe25b243b1df2175d0d788\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}
No param, no code, no length information in any of them.
Client-side consequence. Our harness classifies a context-overflow error by matching message text (prompt is too long, context_length_exceeded, exceeds … maximum context length, etc.). None of these payloads contains any such phrase, so the error is classified as INVALID_REQUEST, not CONTEXT_WINDOW_EXCEEDED. The client therefore neither compacts nor retries: the session is permanently stalled at the point of failure.
Expected
HTTP 200, or an error that names the violated constraint — e.g. prompt_tokens + max_completion_tokens = 646155 > 645120 maximum. Your platform already returns the limit and the requested size elsewhere (cf. #911: 400 This model's maximum context length is 1048576 tokens. However, you requested …).
Evidence — control data, same provider + model id
| Date (UTC) |
Successful calls with prompt > 262,144 |
Max prompt |
422s |
| 2026-09-27 |
60 |
371,246 |
none |
| 2026-09-28 |
311 |
381,743 (last large success 352,480 @11:00:04) |
none |
| 2026-09-29 |
0 |
259,269 (784fcb42 @08:38:32) |
6 in real sessions + 2 probe 422s |
The last observed successful call with prompt > 262,144 was at 2026-09-28 11:00:04 UTC. The first observed 422 was at 2026-09-29 08:38:49 UTC — ≈21.6 hours apart. On 2026-09-29 the largest successful prompt across all sessions was 259,269; session bb836559 was rejected at ≈216,540 with the same route and key.
What we ruled out
- Prompt length alone — 380,012 prompt tokens returned 200, 45% past 262,144.
- Output reservation alone — 249,980 + 384,000 = 633,980 returned 200.
- The output field name —
max_completion_tokens and max_tokens behave identically.
- Request body byte size — 1.63 MB bodies returned 200; a 422 fired on a 1.13 MB body.
- Account, quota, billing — 97 successful calls on the same key after the first failures, 5-hour window usage ≈3–4%, no 429 in history.
- Client-side truncation or retry — we send the conversation as-is; each failing step was one POST, no retry.
INVALID_REQUEST is not in our retry set.
- Request content — no images; 65 tools/52 KB versus 66 tools/53 KB in a successful 09-28 request; flat prefix-cache hit rates across the boundary.
- A different model id — all three days and all eight trace_ids are
commandcode + deepseek/deepseek-v4.1-flash.
- A fixed 262,144-token prompt ceiling — 380,012 prompt tokens returned 200; ≈216,540 returned 422.
- A deterministic size threshold for our client's field set — 09:35:50 UTC ≈212,323 → 422; 09:36:24 UTC 212,336 → 200 (34 seconds later); 09:38:06 UTC 215,824 → 200; 09:39:21 UTC ≈216,540 → 422 (75 seconds later). Same key, model, client configuration, field set and
max_completion_tokens.
- One single
prompt + max_completion_tokens ceiling for all request shapes — the minimal-field shape was accepted at 644,984 (08:57–08:59 UTC) and our client's shape was rejected at ≈600,540 (09:39:21 UTC) on the same day.
Measurement notes. Rejected requests return no usage, so the figures marked ≈ are arithmetic from the preceding successful hop's provider-reported totals (e.g. 215,824 prompt + 716 output → ≈216,540 for the next hop), not direct readings. Every other token figure in this report is a value reported by your API.
Questions for the maintainers
- What is the exact ceiling in force on
deepseek/deepseek-v4.1-flash as of 2026-09-29 09:39 UTC, and what is it measured against — prompt tokens, prompt + max_completion_tokens, or body bytes?
- Is request validation on this route deterministic? Two requests 13 tokens apart (≈212,323 vs 212,336) returned 422 and 200 respectively, 34 seconds apart; two requests 716 tokens apart (215,824 vs ≈216,540) returned 200 and 422 respectively, 75 seconds apart. Same key, same model, same field set, same
max_completion_tokens: 384000.
- Does the validator assert
prompt_tokens + max_completion_tokens ≤ context_window? What is that window for this model — 645,120 (630 Ki)? 646,144 (631 Ki)? Our minimal-field probes bracket it to [644,984, 646,155).
- Did the limit for this route change on 2026-09-29? The same key was accepted with
prompt + max_completion_tokens = 644,984 (minimal fields, 08:57–08:59 UTC) and rejected with ≈600,540 (our field set, 09:39:21 UTC).
- Is
reasoning_effort: "max" accepted on this route? OpenAI's enum is minimal/low/medium/high/none; we send "max". Does the server use the max_completion_tokens value as sent, or a different output reservation, when reasoning_effort is set? Same questions for role:"developer", store:false and stream_options. Probe P-A (prompt + reservation sum only 263,066, our field set) was rejected.
- Trace lookups — which constraint did each of these violate?
f9965636e3da815215348ef4ddf2ae63 — 09:38:49 (prompt ≈262,155, reservation 384,000)
8d2b78e48b922f3e14df82ca87c33073 — 09:39:20 (prompt ≈262,155, reservation 384,000)
b24db3978e79eba528520556aad60aac — 08:57–08:59 (probe EX-1)
691e97aac1847cba8ba8b3a7e011b03f — 08:57–08:59 (probe P-A, reservation 16, sum 263,066)
de1f2c8d39f9ee0b5961a261772e06d4 — 09:35:50 (prompt ≈212,323, reservation 384,000)
789fac7f2c68d85e78c95cbe60a92bdc — 09:39:21 (prompt ≈216,540, reservation 384,000)
5aab8b1aab4f8a81f2c8dbfd1dffcdc4 — 09:39:35 (same)
7d4ff01156fe25b243b1df2175d0d788 — 09:39:43 (same)
- Is
max_completion_tokens: 384000 a valid value on this route? What value should clients send?
- Could this error name the violated limit and the observed size, and return 400 like your other oversized-prompt errors? With the current opaque body, clients cannot distinguish "reduce the reservation and continue" from "the request is malformed", and sessions stay stuck.
Summary
Reproducible. Still failing at the time of writing: 2026-09-29 09:39 UTC.
POST https://api.commandcode.ai/provider/v1/chat/completionswith modeldeepseek/deepseek-v4.1-flashreturns HTTP 422invalid request errorwith no detail — noparam, nocode, no length information, only atrace_id— and the client request is a valid OpenAI chat-completions body.Three measurements on 2026-09-29:
1. Our client's request shape is rejected far below 262,144 (256 Ki) prompt tokens. In session
session-bb836559-…(contextWindow: 1000000,maxTokens: 384000) the endpoint returned 200 for a 215,824-token prompt at 09:38:06 UTC, then 422 for a ≈216,540-token prompt at 09:39:21 UTC — 75 seconds later, same field set, samemax_completion_tokens: 384000. Three consecutive attempts (09:39:21 / 09:39:35 / 09:39:43 UTC) returned 422 and the session stopped responding. 216,540 prompt tokens is 82% of 262,144: the session never reached 256 Ki.2. The same shape also failed at ≈212k, and an identical size succeeded 34 seconds later. 09:35:50 UTC → 422; 09:36:24 UTC → 200 with
prompt_tokens212,336. The two prompts are 13 tokens apart.3. With a minimal field set (
model,messages,max_completion_tokensonly) the accepted/rejected boundary tracksprompt_tokens + max_completion_tokens. 200 up to 644,984; 422 from 646,155. A 380,012-token prompt with a reservation of 16 returned 200. Direct probes, same route, same key, 08:57–08:59 UTC today.Impact. With
max_completion_tokens: 384000the largest prompt we ever got through was 260,984 tokens (probe) and 259,269 tokens (real session). Two real sessions died at that point and did not recover:session-784fcb42-…at 259,269 (08:38 UTC) andsession-bb836559-…at 215,824 (09:38 UTC). Your documentation advertises up to 1M context for this route. Clients cannot recover, because the error body names no limit (see Questions 8).Timeline — 2026-09-29, all times UTC, all
commandcode+deepseek/deepseek-v4.1-flash, same API keyusage)f9965636e3da815215348ef4ddf2ae638d2b78e48b922f3e14df82ca87c33073b24db3978e79eba528520556aad60aac691e97aac1847cba8ba8b3a7e011b03fde1f2c8d39f9ee0b5961a261772e06d4789fac7f2c68d85e78c95cbe60a92bdc5aab8b1aab4f8a81f2c8dbfd1dffcdc47d4ff01156fe25b243b1df2175d0d788≈= arithmetic from the preceding successful hop's provider-reported totals (rejected requests return nousage; see Measurement notes).Probe results with a minimal field set:
Environment
POST https://api.commandcode.ai/provider/v1/chat/completions(no proxy in the path)deepseek/deepseek-v4.1-flashcommandcode, apiopenai-completions, declared contextWindow 1,000,000 / maxTokens 384,000stream:true,stream_options:{include_usage:true},store:false,reasoning_effort:"max",role:"developer"for the system turn,tools:[…]Steps to reproduce — we executed this and it reproduces
→ HTTP 422, body ≈1.13 MB, ≈263,050 prompt tokens,
content-type: application/json(nottext/event-stream), no SSE events, 165-byte error body.Contrast runs, all 200:
max_completion_tokens384,000 → 16, no other field changes → 200max_tokensinstead → 200Token counting on your side is deterministic, so prompt sizes are exact:
prompt_tokens = 0.2333333 × filler_chars + 38. Calibrated on two points, then used to predict three further requests with 0-token error (249,980 / 270,014 / 380,012 all exact).Observed
Real session
session-784fcb42-…, 2026-09-29. 08:38:32 UTC success: yourusagereports input 8,773 + cache read 250,496 = 259,269 prompt tokens, output 2,886 (total_tokens262,155). 08:38:49 UTC next hop, prompt ≈ 262,155 → 422. 08:39:20 UTC the user typed "continue" → 422 again. Every later request in that session returned 422; it did not recover. Other sessions on the same key/provider kept working (97 successful calls after 08:38:49, max prompt 166,677).Real session
session-bb836559-…, 2026-09-29. Table above: 200 at 215,824 prompt tokens (09:38:06), 422 at ≈216,540 (09:39:21 / 09:39:35 / 09:39:43). The session stopped responding at 09:39:43 and did not recover. It had also been rejected once at ≈212,323 (09:35:50) and accepted 34 seconds later at 212,336 (09:36:24).Raw 422 payloads — 8 responses, 8 distinct trace_ids:
No
param, nocode, no length information in any of them.Client-side consequence. Our harness classifies a context-overflow error by matching message text (
prompt is too long,context_length_exceeded,exceeds … maximum context length, etc.). None of these payloads contains any such phrase, so the error is classified asINVALID_REQUEST, notCONTEXT_WINDOW_EXCEEDED. The client therefore neither compacts nor retries: the session is permanently stalled at the point of failure.Expected
HTTP 200, or an error that names the violated constraint — e.g.
prompt_tokens + max_completion_tokens = 646155 > 645120 maximum. Your platform already returns the limit and the requested size elsewhere (cf. #911:400 This model's maximum context length is 1048576 tokens. However, you requested …).Evidence — control data, same provider + model id
The last observed successful call with prompt > 262,144 was at 2026-09-28 11:00:04 UTC. The first observed 422 was at 2026-09-29 08:38:49 UTC — ≈21.6 hours apart. On 2026-09-29 the largest successful prompt across all sessions was 259,269; session
bb836559was rejected at ≈216,540 with the same route and key.What we ruled out
max_completion_tokensandmax_tokensbehave identically.INVALID_REQUESTis not in our retry set.commandcode+deepseek/deepseek-v4.1-flash.max_completion_tokens.prompt + max_completion_tokensceiling for all request shapes — the minimal-field shape was accepted at 644,984 (08:57–08:59 UTC) and our client's shape was rejected at ≈600,540 (09:39:21 UTC) on the same day.Measurement notes. Rejected requests return no
usage, so the figures marked≈are arithmetic from the preceding successful hop's provider-reported totals (e.g. 215,824 prompt + 716 output → ≈216,540 for the next hop), not direct readings. Every other token figure in this report is a value reported by your API.Questions for the maintainers
deepseek/deepseek-v4.1-flashas of 2026-09-29 09:39 UTC, and what is it measured against — prompt tokens, prompt +max_completion_tokens, or body bytes?max_completion_tokens: 384000.prompt_tokens + max_completion_tokens ≤ context_window? What is that window for this model — 645,120 (630 Ki)? 646,144 (631 Ki)? Our minimal-field probes bracket it to [644,984, 646,155).prompt + max_completion_tokens = 644,984(minimal fields, 08:57–08:59 UTC) and rejected with ≈600,540 (our field set, 09:39:21 UTC).reasoning_effort: "max"accepted on this route? OpenAI's enum isminimal/low/medium/high/none; we send"max". Does the server use themax_completion_tokensvalue as sent, or a different output reservation, whenreasoning_effortis set? Same questions forrole:"developer",store:falseandstream_options. Probe P-A (prompt + reservationsum only 263,066, our field set) was rejected.f9965636e3da815215348ef4ddf2ae63— 09:38:49 (prompt ≈262,155, reservation 384,000)8d2b78e48b922f3e14df82ca87c33073— 09:39:20 (prompt ≈262,155, reservation 384,000)b24db3978e79eba528520556aad60aac— 08:57–08:59 (probe EX-1)691e97aac1847cba8ba8b3a7e011b03f— 08:57–08:59 (probe P-A, reservation 16, sum 263,066)de1f2c8d39f9ee0b5961a261772e06d4— 09:35:50 (prompt ≈212,323, reservation 384,000)789fac7f2c68d85e78c95cbe60a92bdc— 09:39:21 (prompt ≈216,540, reservation 384,000)5aab8b1aab4f8a81f2c8dbfd1dffcdc4— 09:39:35 (same)7d4ff01156fe25b243b1df2175d0d788— 09:39:43 (same)max_completion_tokens: 384000a valid value on this route? What value should clients send?