Skip to content

Repository files navigation

Konnect Cost Reporter for Finout

Reference implementation for exporting internal usage cost from Kong Konnect Metering & Billing to a Finout Custom Cost Center.

The repository builds one executable. Each invocation:

  1. reads configuration from environment variables
  2. queries cost-enabled Konnect features for completed UTC days
  3. converts the result to one Finout CSV per day
  4. writes those files to S3 or an S3-compatible object store
  5. exits

Production scheduling is intentionally outside the process. The included Helm chart runs the executable as a Kubernetes CronJob.

Architecture

flowchart LR
    Gateway["Kong Gateway<br/>Metering & Billing plugin"]
    Events["Usage events<br/>kong.api_request / kong.llm_request"]
    Meters["Konnect meters<br/>aggregate usage"]
    Features["Konnect features<br/>filters + unit cost"]
    Query["Cost Analytics API<br/>daily cost rows"]
    Reporter["finout-cost-reporter"]
    S3["S3 daily CSVs"]
    Finout["Finout<br/>Custom Cost Center"]

    Gateway --> Events --> Meters --> Features --> Query --> Reporter --> S3 --> Finout
Loading

Konnect performs the cost calculation. The reporter does not maintain an LLM price table or multiply usage locally. It exports Cost Analytics results.

How Metered Cost Works

The Metering & Billing plugin emits immutable usage events from Kong Gateway. The event subject identifies the customer or tenant responsible for usage, while event properties such as service, route, provider, model, and token type provide allocation dimensions.

A meter selects an event type and aggregates a value:

  • API requests use count over kong.api_request.
  • LLM usage uses sum over the event's tokens property for kong.llm_request.

A feature turns metered usage into something that can be priced or governed. Its meter filters select a slice of usage. Its unit cost describes the provider's internal cost:

  • API requests use a manual per-request cost, $0.01 in the demo.
  • LLM token features use provider, model, and token-type filters, then resolve per-token cost from Konnect's automatically maintained LLM cost database.

Unit cost is internal cost, not customer price. Customer-facing prices belong in Product Catalog rate cards. See Product Catalog and unit cost.

An LLM unit cost is resolved from three things: provider, model, and token type. The AI Gateway emits two kong.llm_request event types in the type field: request (prompt / input tokens) and response (completion / output tokens). The cost database prices them per token (input_per_token / output_per_token). For token_type, request and response are accepted as aliases of input / output. The gateway does not emit a cache or reasoning token type, so only input and output are modeled. A feature can supply each value in either of two ways:

  • Statically. Pin the feature to one provider, model, and token type (the unit_cost.llm provider / model / token_type fields). This repo's demo uses this approach: one feature per provider × model × event type.
  • From a meter dimension. Let the value come from the event via a meter group-by property (provider_property / model_property / token_type_property), so a single feature prices many providers/models at once.

Either way, the meter must declare provider, model, and type as group-by dimensions. This is not for pricing (the demo's prices come from the static values). It is because the feature's meter filters narrow usage on exactly those dimensions (provider = openai, model = …, type = request). All three are part of Konnect's standard LLM Tokens meter template, so a meter created from that template already has them. The demo meter includes them explicitly for the same reason. Request and response tokens are priced separately because providers charge different rates for prompt (input) and completion (output) tokens.

LLM cost database: freshness and timing

The LLM cost database is maintained by Kong, not by this reporter. A Konnect background job (llm-cost sync) refreshes prices from upstream pricing sources every 6 hours. That is a refresh interval, not a freshness guarantee: the upstream sources the job reads can themselves lag a provider's actual change or a new model's release, so a price can stay stale for longer than 6 hours even though the sync runs on schedule. Treat the database as eventually-consistent with the real world. It can therefore lag real-world pricing in two ways:

  • A provider changes a price. Until Kong's database picks up the new rate, cost for that model is computed at the previous rate. When the new price lands it is recorded as a new effective period (see below). Already-exported days keep the price that was in effect at the time and are only restated if you re-run that day after the database updates.
  • A new model ships. A model with no entry in the database yet is unpriced: the cost query returns a row with cost = null and a detail explaining the gap. This reporter skips such rows with a warning rather than emitting 0.00, so missing prices never silently understate spend. The affected day is recorded as pending in the state marker and re-exported automatically on every subsequent run. Its CSV is overwritten with the full cost as soon as Kong adds the model (or you add an override). No manual intervention is needed. See Delivery guarantees & catch-up.

Cost is computed at the price effective at the time of usage, not today's price. Each price carries effective_from / effective_to timestamps, and Konnect picks the record covering each usage day. So a backfill (QUERY_DURATION = P7D) prices each day with that day's rate, and re-running an old day after a price change can legitimately change its cost. Because object keys are deterministic per day, such a re-run overwrites the day and Finout re-ingests the corrected figure.

Closing the gap yourself. For a model the database lacks, or a rate you've negotiated below list price, create a per-org price override (provider + model + per-token pricing). Overrides take precedence over Kong's system prices, so a re-run immediately resolves cost for the affected rows. The override is managed in Konnect (the SDK exposes llmCost.listOverrides / createOverride). This reporter only reads the resulting cost. See Cost Analytics for the database, effective dating, and overrides.

Demo Resources

The exporter reads existing meters and features. It never creates them. If you want a working set to try it against, an optional helper in src/demo.ts provisions sample resources. Run it explicitly, and only when you want it:

pnpm demo:setup

This writes to your Konnect organization (it needs a read/write token), so it is a deliberate, one-time action. It is not part of pnpm build, the exporter, or the Helm CronJob, and nothing runs it automatically.

It creates or updates, idempotently:

  • kong_konnect_api_request, counting all API Gateway request events
  • kong_konnect_llm_tokens, summing AI Gateway token events
  • kong_api_requests, with the manual API_REQUEST_UNIT_COST
  • two LLM features per configured demo model: one for request (input) tokens and one for response (output) tokens

By default the LLM models are:

  • openai:chatgpt-4o-latest
  • anthropic:claude-3-5-sonnet

Each LLM feature has exact provider, model, and type meter filters and an LLM unit cost pointing at the same provider/model/token type in Konnect's cost database. Change DEMO_LLM_MODELS to a comma-separated list of provider:model pairs.

It refuses to silently replace incompatible existing meter or feature definitions.

Local Setup

Requirements:

  • devenv
  • direnv is optional but recommended
  • a Konnect token with Metering & Billing read access
cp .env.example .env
# Set API_TOKEN and API_BASE_URL for the Konnect region.

direnv allow
pnpm install
devenv up -d
pnpm build
node --env-file=.env dist/index.js

Garage listens on http://127.0.0.1:3900, creates the finout-cost-reports bucket, and grants the development key from .env.example access to it. Its local UI is available at http://127.0.0.1:3919.

This flow creates no meters or features. The reporter only reads from Konnect. If you need sample meters and features to read from, run the optional demo setup described in Demo Resources. It writes to your Konnect organization, so it needs a read/write token and is never run as part of normal operation.

For a query-only check, set DRY_RUN=true. The reporter still calls Konnect and builds CSVs, but it does not write objects.

Delivery guarantees & catch-up

The exporter is safe to run at any time, whether on schedule, manually, or after a gap. It converges to every complete UTC day present in S3 exactly once.

No duplicates. Each day is written to a deterministic key, s3://<bucket>/<prefix>/YYYY-MM-DD.csv, with an overwriting PutObject. Running the same day again replaces that object. It never creates a second file. Finout re-ingests a changed file and replaces that day's data.

No gaps, even when runs are delayed or skipped. The exporter records its progress in a small marker object, s3://<bucket>/<prefix>/_state.json, holding the last UTC day processed and a list of pendingDates (see below). Each run:

  1. reads the marker (absent on the first run → it bootstraps using QUERY_DURATION)
  2. exports every complete day from the day after the marker up to yesterday, so any missed days are filled in
  3. advances the marker, but only after all that run's days uploaded successfully, and only ever forward

Because the window starts at the oldest un-exported day, a delayed or skipped run is caught up by the next run, however long the gap. A failed run doesn't advance the marker, so the next run safely retries the same days (idempotent, by the deterministic keys).

Single writer. The marker is read-modify-written without locking, so it assumes one run at a time. The Helm CronJob enforces this with concurrencyPolicy: Forbid. If that is ever violated (e.g. a manual kubectl create job racing the cron), a late-finishing older run could momentarily set the marker back. The very next run re-exports the affected days and re-advances, so the worst case is repeated work, never lost or duplicated data. Don't run the job concurrently against the same bucket/prefix.

Incomplete pricing self-heals. When a day exports but some usage has no price yet (a new model before Kong's cost-DB sync, or one needing an override), that day's CSV is written with the rows that did price and the date is added to pendingDates. The marker still advances past it, so a never-priced day can never stall the export. Every later run re-queries the pending days and overwrites their CSVs once pricing resolves, dropping them from the set. This is fully automatic (a warn log each run, no alert). See LLM cost database.

Blast-radius cap. A single run exports at most MAX_BACKFILL_DAYS days (default 31) in the forward window. Pending days are re-queried in addition, each as its own targeted query so a stuck day can't widen the window. A longer outage is drained oldest-first across successive runs until it catches up, and each run that leaves a backlog logs a warning. A consequence worth knowing: while draining a long backlog, the most recent days are exported last. Raise MAX_BACKFILL_DAYS to catch up faster in one run (Konnect accepts multi-month query windows), or trigger the job repeatedly.

To force a re-export of a day or a range, delete or rewind _state.json (set lastExportedDate to the day before the earliest you want re-sent) and run again. The deterministic keys mean the affected days are overwritten in place.

Configuration

Variable Required Default Purpose
API_BASE_URL yes - Regional Konnect API URL, for example https://us.api.konghq.com/v3
API_TOKEN yes - Konnect PAT or system-account token
QUERY_DURATION no P1D First-run only bootstrap window (e.g. P1D, P7D). After the first run the window is driven by the S3 state marker. See Delivery guarantees & catch-up
MAX_BACKFILL_DAYS no 31 Max UTC days one run may export. A longer backlog drains oldest-first across runs
FEATURE_KEYS no all priced features Comma-separated feature keys or IDs
GROUP_BY_DIMENSIONS no subject Extra meter dimensions to group cost by and append as CSV columns. subject is always included (customer_id is not groupable and is ignored if listed)
S3_ENDPOINT_URL no AWS endpoint Garage, MinIO, or another custom endpoint
S3_REGION no us-east-1 S3 signing/bucket region
S3_BUCKET yes - Destination bucket
S3_PREFIX no empty Object key prefix
S3_FORCE_PATH_STYLE no false Required by Garage and many compatible stores
S3_ACCESS_KEY_ID no AWS SDK chain Static object-store access key
S3_SECRET_ACCESS_KEY no AWS SDK chain Static object-store secret
LOG_LEVEL no info debug, info, warn, or error
DRY_RUN no false Skip S3 writes
API_REQUEST_UNIT_COST demo only 0.01 Internal USD cost per API request
DEMO_LLM_MODELS demo only two example models Comma-separated provider:model pairs

Standard AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY are also accepted. When no static credentials are supplied, the AWS SDK credential chain supports profiles, ECS credentials, EC2 roles, EKS IRSA, and EKS Pod Identity.

Finout Output

Finout requires one CSV file per day with UsageDate and USD Cost. Every other column is a cost-allocation dimension. This reporter writes FeatureName, FeatureKey, and the usage-attribution Subject, plus a Customer column resolved from each subject (see "Cost attribution" below). Any extra meter group-by dimensions you enable (see GROUP_BY_DIMENSIONS) are appended as further columns:

"UsageDate","Cost","FeatureName","FeatureKey","Subject","Customer"
"2026-06-05","0.015","OpenAI chatgpt-4o-latest input tokens","kong_llm_openai_chatgpt_4o_latest_input","team-alpha","Acme Corp"

With GROUP_BY_DIMENSIONS=model,provider the same row gains trailing model and provider columns. Customer is omitted entirely on days where no subject maps to a customer.

Cost attribution

Cost is broken down by subject, the dimension that identifies who incurred the usage. The cost query groups by subject only. It deliberately does not group by customer_id, because the Cost Analytics endpoint rejects that without a per-customer filter (customer filter is required with customer_id group by) and so cannot drive a single all-customers export.

The Customer column is filled in afterwards. A Konnect customer declares the subjects attributed to it (usage_attribution.subject_keys). The reporter lists all customers, inverts that into a subject → customer map, and labels each row by its subject. So when several subjects belong to one customer, each subject is still its own row. Subjects are never summed together. A subject with no customer gets a blank Customer.

Objects use deterministic keys:

s3://<bucket>/<prefix>/YYYY-MM-DD.csv

Re-running a day replaces that day's object. Finout ingests newly created or updated files and replaces previously ingested data when an older file changes. This makes retries and P7D backfills idempotent.

Rows whose LLM price cannot be resolved are skipped with a warning instead of being reported as zero cost. The affected day is marked pending and re-exported automatically once pricing resolves (see LLM cost database: freshness and timing). Non-USD cost fails the run because Finout's custom cost contract expects USD. A completed day with no rows is still uploaded as a header-only CSV so a retry can clear stale data from an earlier export.

To connect the bucket, grant Finout read access to the configured prefix and follow its Custom Cost Center onboarding process. Finout currently asks customers to contact support with the S3-hosted sample CSV.

Production

Build the image:

docker build -t ghcr.io/openmeterio/finout-cost-reporter:0.1.0 .

Install the CronJob with chart-managed secrets:

helm upgrade --install finout-cost-reporter \
  deploy/helm/finout-cost-reporter \
  --namespace finout --create-namespace \
  --set secret.apiToken="$API_TOKEN" \
  --set secret.s3AccessKeyId="$S3_ACCESS_KEY_ID" \
  --set secret.s3SecretAccessKey="$S3_SECRET_ACCESS_KEY" \
  --set config.s3.bucket="my-finout-cost-bucket"

For AWS workload identity, omit the S3 keys and annotate the service account:

secret:
  apiToken: kpat_REPLACE_ME

serviceAccount:
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/finout-reporter
  automountServiceAccountToken: true

The reporter's S3 identity needs s3:PutObject (write each day's CSV and the state marker) and s3:GetObject (read the state marker to resume) on arn:aws:s3:::<bucket>/<prefix>/*. The GetObject permission is required. A write-only policy makes every run re-bootstrap and can leave gaps. (Finout's own read access to the bucket is separate. See below.)

For production secret management, set secret.create=false and secret.existingSecret to a Secret containing API_TOKEN. Static S3 keys are optional.

The default schedule is 01:00 UTC. P1D always queries the previous complete UTC day, regardless of the actual start time. concurrencyPolicy: Forbid prevents overlapping exports.

Development

pnpm check        # format check + lint + typecheck + test (the CI gate)
pnpm fmt          # auto-format with oxfmt
pnpm lint:fix     # auto-fix lint issues with oxlint
pnpm typecheck
pnpm test
helm lint deploy/helm/finout-cost-reporter \
  -f deploy/helm/finout-cost-reporter/ci/test-values.yaml
helm template test deploy/helm/finout-cost-reporter \
  -f deploy/helm/finout-cost-reporter/ci/test-values.yaml

Code is formatted with oxfmt and linted with oxlint. One enforced rule: if/else/loop bodies are always braced (curly).

API shapes are provided by @openmeter/client, generated from the Konnect Metering & Billing OpenAPI specification. Garage configuration follows the devenv Garage service.

About

Reference implementation for exporting metered feature cost from Kong Konnect Metering & Billing to a Finout Custom Cost Center.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages