Skip to content

[Regression] Fix for #1878 only covers 'openai' provider — custom: OpenAI-compatible endpoints still send max_tokens to gpt-5 family #2415

Description

@Popcorncandy09

The fix in commit resolving #1878 (rename max_tokens → max_completion_tokens for gpt-5 family) is gated on provider === "openai" and does not apply when an OpenAI-compatible upstream is registered as a custom: provider. This means a custom: provider fronting an Azure OpenAI / Azure Foundry / vLLM / LiteLLM / OpenRouter-style deployment that serves gpt-5, gpt-5-mini, gpt-5-nano, or any o-series model still receives max_tokens in the outbound body and is rejected by the upstream with HTTP 400 unsupported_parameter.

This is the same observable symptom as #1878, so users with hybrid deployments (subscription OpenAI for some models, Azure Foundry for others) will see the same 400 they reported before the fix.

Environment

  • Manifest version: latest (image digest sha256:26d2a86d3dcc693512581f6aa3b5a9ef8b3bf9c31944fda9e638f722b47821f3, pulled 2026-07-05)
  • Self-hosted (MANIFEST_MODE=selfhosted)
  • Affected providers: any custom: provider where the underlying endpoint is OpenAI-API-compatible and serves gpt-5 / gpt-5-* / o1 / o3 / o4 models
  • Confirmed reproducing with Azure Foundry endpoint: https://dangpt-az.services.ai.azure.com/openai/v1/chat/completions serving gpt-5.4-mini

Reproduction steps

  1. Register a custom provider in Manifest with base_url pointing to an OpenAI-compatible endpoint serving a gpt-5 family model (e.g. Azure Foundry, vLLM, LiteLLM proxy).

  2. Configure the auto routing tier (or any direct model request) to pick that custom provider's gpt-5-mini (or gpt-5, gpt-5-nano, o1, o3-mini, o4-mini, etc.).

  3. Send a chat completion:

    curl -X POST http://manifest:2099/v1/chat/completions
    -H "Authorization: Bearer ***"
    -H "Content-Type: application/json"
    -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "PONG"}],
    "max_tokens": 50
    }'

  4. Observe the response: HTTP 400 with body Bad request to upstream provider (Manifest wraps the upstream's 400, so the inner error is only visible in container logs).

  5. Check the Manifest container log:

    [ProviderClient] Forwarding to custom: https://dangpt-az.services.ai.azure.com/openai/v1/chat/completions
    WARN [ProxyResponseHandler] Upstream error 400: ... body={
    "error": {
    "message": "Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.",
    "type": "invalid_request_error",
    "param": "max_tokens",
    "code": "unsupported_parameter"
    }
    }

  6. Compare: sending the same request to Manifest with the same model but configured as a direct openai provider instead of custom: returns 429 (rate limit) — proving the param validation passed in that path, confirming the fix is active only for provider === "openai".

Expected behavior

The fix from #1878 should apply to any OpenAI-compatible upstream that requires max_completion_tokens, not just direct OpenAI. The provider kind (openai vs custom:) is orthogonal to which parameter the upstream model accepts.

Actual behavior

When the upstream is registered as a custom: provider, Manifest still sends max_tokens to gpt-5 family and o-series models. The upstream rejects with HTTP 400 unsupported_parameter, Manifest wraps the error as Bad request to upstream provider, the caller's request fails, and the fallback tier (if any) takes the cost of the failed call.

Suggested fix

In packages/backend/src/common/constants/openai-models.ts (or wherever the rename lives after #1878), extend OPENAI_USES_MAX_COMPLETION_TOKENS_RE matching to also cover the custom: provider path. Concretely, the rename should fire when any of the following are true:

  • provider is openai AND model matches the gpt-5 / o-series regex
  • provider is custom: AND the underlying custom_providers.models array contains a model name matching the gpt-5 / o-series regex

The matching should use the resolved downstream model name (after Manifest's router picks it), not the client's request model name, because the client often sends auto or a tier alias and Manifest decides the real model.

Related


Or can we have option to connect direct to Azure Foundry

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions