Reference implementation for exporting internal usage cost from Kong Konnect Metering & Billing to a Finout Custom Cost Center.
The repository builds one executable. Each invocation:
- reads configuration from environment variables
- queries cost-enabled Konnect features for completed UTC days
- converts the result to one Finout CSV per day
- writes those files to S3 or an S3-compatible object store
- exits
Production scheduling is intentionally outside the process. The included Helm chart runs the executable as a Kubernetes CronJob.
flowchart LR
Gateway["Kong Gateway<br/>Metering & Billing plugin"]
Events["Usage events<br/>kong.api_request / kong.llm_request"]
Meters["Konnect meters<br/>aggregate usage"]
Features["Konnect features<br/>filters + unit cost"]
Query["Cost Analytics API<br/>daily cost rows"]
Reporter["finout-cost-reporter"]
S3["S3 daily CSVs"]
Finout["Finout<br/>Custom Cost Center"]
Gateway --> Events --> Meters --> Features --> Query --> Reporter --> S3 --> Finout
Konnect performs the cost calculation. The reporter does not maintain an LLM price table or multiply usage locally. It exports Cost Analytics results.
The Metering & Billing plugin emits immutable usage events from Kong Gateway. The event subject identifies the customer or tenant responsible for usage, while event properties such as service, route, provider, model, and token type provide allocation dimensions.
A meter selects an event type and aggregates a value:
- API requests use
countoverkong.api_request. - LLM usage uses
sumover the event'stokensproperty forkong.llm_request.
A feature turns metered usage into something that can be priced or governed. Its meter filters select a slice of usage. Its unit cost describes the provider's internal cost:
- API requests use a manual per-request cost,
$0.01in the demo. - LLM token features use provider, model, and token-type filters, then resolve per-token cost from Konnect's automatically maintained LLM cost database.
Unit cost is internal cost, not customer price. Customer-facing prices belong in Product Catalog rate cards. See Product Catalog and unit cost.
An LLM unit cost is resolved from three things: provider, model, and
token type. The AI Gateway emits two kong.llm_request event types in the
type field: request (prompt / input tokens) and response
(completion / output tokens). The cost database prices them per token
(input_per_token / output_per_token). For token_type, request and
response are accepted as aliases of input / output. The gateway does
not emit a cache or reasoning token type, so only input and output are
modeled.
A feature can supply each value in either of two ways:
- Statically. Pin the feature to one provider, model, and token type (the
unit_cost.llmprovider/model/token_typefields). This repo's demo uses this approach: one feature perprovider × model ×event type. - From a meter dimension. Let the value come from the event via a meter
group-by property (
provider_property/model_property/token_type_property), so a single feature prices many providers/models at once.
Either way, the meter must declare provider, model, and type as
group-by dimensions. This is not for pricing (the demo's prices come from the
static values). It is because the feature's meter filters narrow usage on
exactly those dimensions (provider = openai, model = …, type = request).
All three are part of Konnect's standard LLM Tokens meter template, so a
meter created from that template already has them. The demo meter includes them
explicitly for the same reason. Request and response tokens are priced
separately because providers charge different rates for prompt (input) and
completion (output) tokens.
The LLM cost database is maintained by Kong, not by this reporter. A Konnect
background job (llm-cost sync) refreshes prices from upstream pricing sources
every 6 hours. That is a refresh interval, not a freshness guarantee:
the upstream sources the job reads can themselves lag a provider's actual change
or a new model's release, so a price can stay stale for longer than 6 hours even
though the sync runs on schedule. Treat the database as eventually-consistent
with the real world. It can therefore lag real-world pricing in two ways:
- A provider changes a price. Until Kong's database picks up the new rate, cost for that model is computed at the previous rate. When the new price lands it is recorded as a new effective period (see below). Already-exported days keep the price that was in effect at the time and are only restated if you re-run that day after the database updates.
- A new model ships. A model with no entry in the database yet is
unpriced: the cost query returns a row with
cost = nulland adetailexplaining the gap. This reporter skips such rows with a warning rather than emitting0.00, so missing prices never silently understate spend. The affected day is recorded as pending in the state marker and re-exported automatically on every subsequent run. Its CSV is overwritten with the full cost as soon as Kong adds the model (or you add an override). No manual intervention is needed. See Delivery guarantees & catch-up.
Cost is computed at the price effective at the time of usage, not today's
price. Each price carries effective_from / effective_to timestamps, and
Konnect picks the record covering each usage day. So a backfill (QUERY_DURATION = P7D) prices each day with that day's rate, and re-running an old day after a
price change can legitimately change its cost. Because object keys are
deterministic per day, such a re-run overwrites the day and Finout re-ingests the
corrected figure.
Closing the gap yourself. For a model the database lacks, or a rate you've
negotiated below list price, create a per-org price override (provider +
model + per-token pricing). Overrides take precedence over Kong's system prices,
so a re-run immediately resolves cost for the affected rows. The override is
managed in Konnect (the SDK exposes llmCost.listOverrides /
createOverride). This reporter only reads the resulting cost. See
Cost Analytics
for the database, effective dating, and overrides.
The exporter reads existing meters and features. It never creates them. If you
want a working set to try it against, an optional helper in src/demo.ts
provisions sample resources. Run it explicitly, and only when you want it:
pnpm demo:setupThis writes to your Konnect organization (it needs a read/write token), so
it is a deliberate, one-time action. It is not part of pnpm build, the
exporter, or the Helm CronJob, and nothing runs it automatically.
It creates or updates, idempotently:
kong_konnect_api_request, counting all API Gateway request eventskong_konnect_llm_tokens, summing AI Gateway token eventskong_api_requests, with the manualAPI_REQUEST_UNIT_COST- two LLM features per configured demo model: one for
request(input) tokens and one forresponse(output) tokens
By default the LLM models are:
openai:chatgpt-4o-latestanthropic:claude-3-5-sonnet
Each LLM feature has exact provider, model, and type meter filters and an
LLM unit cost pointing at the same provider/model/token type in Konnect's cost
database. Change DEMO_LLM_MODELS to a comma-separated list of
provider:model pairs.
It refuses to silently replace incompatible existing meter or feature definitions.
Requirements:
- devenv
- direnv is optional but recommended
- a Konnect token with Metering & Billing read access
cp .env.example .env
# Set API_TOKEN and API_BASE_URL for the Konnect region.
direnv allow
pnpm install
devenv up -d
pnpm build
node --env-file=.env dist/index.jsGarage listens on http://127.0.0.1:3900, creates the
finout-cost-reports bucket, and grants the development key from
.env.example access to it. Its local UI is available at
http://127.0.0.1:3919.
This flow creates no meters or features. The reporter only reads from Konnect. If you need sample meters and features to read from, run the optional demo setup described in Demo Resources. It writes to your Konnect organization, so it needs a read/write token and is never run as part of normal operation.
For a query-only check, set DRY_RUN=true. The reporter still calls Konnect and
builds CSVs, but it does not write objects.
The exporter is safe to run at any time, whether on schedule, manually, or after a gap. It converges to every complete UTC day present in S3 exactly once.
No duplicates. Each day is written to a deterministic key,
s3://<bucket>/<prefix>/YYYY-MM-DD.csv, with an overwriting PutObject. Running
the same day again replaces that object. It never creates a second file. Finout
re-ingests a changed file and replaces that day's data.
No gaps, even when runs are delayed or skipped. The exporter records its
progress in a small marker object, s3://<bucket>/<prefix>/_state.json, holding
the last UTC day processed and a list of pendingDates (see below). Each
run:
- reads the marker (absent on the first run → it bootstraps using
QUERY_DURATION) - exports every complete day from the day after the marker up to yesterday, so any missed days are filled in
- advances the marker, but only after all that run's days uploaded successfully, and only ever forward
Because the window starts at the oldest un-exported day, a delayed or skipped run is caught up by the next run, however long the gap. A failed run doesn't advance the marker, so the next run safely retries the same days (idempotent, by the deterministic keys).
Single writer. The marker is read-modify-written without locking, so it
assumes one run at a time. The Helm CronJob enforces this with
concurrencyPolicy: Forbid. If that is ever violated (e.g. a manual
kubectl create job racing the cron), a late-finishing older run could
momentarily set the marker back. The very next run re-exports the affected days
and re-advances, so the worst case is repeated work, never lost or duplicated
data. Don't run the job concurrently against the same bucket/prefix.
Incomplete pricing self-heals. When a day exports but some usage has no price
yet (a new model before Kong's cost-DB sync, or one needing an override), that
day's CSV is written with the rows that did price and the date is added to
pendingDates. The marker still advances past it, so a never-priced day can
never stall the export. Every later run re-queries the pending days and
overwrites their CSVs once pricing resolves, dropping them from the set. This is
fully automatic (a warn log each run, no alert). See
LLM cost database.
Blast-radius cap. A single run exports at most MAX_BACKFILL_DAYS days
(default 31) in the forward window. Pending days are re-queried in addition,
each as its own targeted query so a stuck day can't widen the window. A longer
outage is drained oldest-first across successive runs until it catches up,
and each run that leaves a backlog logs a warning. A consequence worth knowing:
while draining a long backlog, the most recent days are exported last. Raise
MAX_BACKFILL_DAYS to catch up faster in one run (Konnect accepts multi-month
query windows), or trigger the job repeatedly.
To force a re-export of a day or a range, delete or rewind _state.json (set
lastExportedDate to the day before the earliest you want re-sent) and run
again. The deterministic keys mean the affected days are overwritten in place.
| Variable | Required | Default | Purpose |
|---|---|---|---|
API_BASE_URL |
yes | - | Regional Konnect API URL, for example https://us.api.konghq.com/v3 |
API_TOKEN |
yes | - | Konnect PAT or system-account token |
QUERY_DURATION |
no | P1D |
First-run only bootstrap window (e.g. P1D, P7D). After the first run the window is driven by the S3 state marker. See Delivery guarantees & catch-up |
MAX_BACKFILL_DAYS |
no | 31 |
Max UTC days one run may export. A longer backlog drains oldest-first across runs |
FEATURE_KEYS |
no | all priced features | Comma-separated feature keys or IDs |
GROUP_BY_DIMENSIONS |
no | subject |
Extra meter dimensions to group cost by and append as CSV columns. subject is always included (customer_id is not groupable and is ignored if listed) |
S3_ENDPOINT_URL |
no | AWS endpoint | Garage, MinIO, or another custom endpoint |
S3_REGION |
no | us-east-1 |
S3 signing/bucket region |
S3_BUCKET |
yes | - | Destination bucket |
S3_PREFIX |
no | empty | Object key prefix |
S3_FORCE_PATH_STYLE |
no | false |
Required by Garage and many compatible stores |
S3_ACCESS_KEY_ID |
no | AWS SDK chain | Static object-store access key |
S3_SECRET_ACCESS_KEY |
no | AWS SDK chain | Static object-store secret |
LOG_LEVEL |
no | info |
debug, info, warn, or error |
DRY_RUN |
no | false |
Skip S3 writes |
API_REQUEST_UNIT_COST |
demo only | 0.01 |
Internal USD cost per API request |
DEMO_LLM_MODELS |
demo only | two example models | Comma-separated provider:model pairs |
Standard AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY are also accepted.
When no static credentials are supplied, the AWS SDK credential chain supports
profiles, ECS credentials, EC2 roles, EKS IRSA, and EKS Pod Identity.
Finout requires one CSV file per day with UsageDate and USD Cost. Every
other column is a cost-allocation dimension. This reporter writes
FeatureName, FeatureKey, and the usage-attribution Subject, plus a
Customer column resolved from each subject (see "Cost attribution" below). Any
extra meter group-by dimensions you enable (see GROUP_BY_DIMENSIONS) are
appended as further columns:
"UsageDate","Cost","FeatureName","FeatureKey","Subject","Customer"
"2026-06-05","0.015","OpenAI chatgpt-4o-latest input tokens","kong_llm_openai_chatgpt_4o_latest_input","team-alpha","Acme Corp"With GROUP_BY_DIMENSIONS=model,provider the same row gains trailing model
and provider columns. Customer is omitted entirely on days where no subject
maps to a customer.
Cost is broken down by subject, the dimension that identifies who
incurred the usage. The cost query groups by subject only. It deliberately does
not group by customer_id, because the Cost Analytics endpoint rejects that
without a per-customer filter (customer filter is required with customer_id group by) and so cannot drive a single all-customers export.
The Customer column is filled in afterwards. A Konnect customer declares the
subjects attributed to it (usage_attribution.subject_keys). The reporter lists
all customers, inverts that into a subject → customer map, and labels each row
by its subject. So when several subjects belong to one customer, each subject
is still its own row. Subjects are never summed together. A subject with no
customer gets a blank Customer.
Objects use deterministic keys:
s3://<bucket>/<prefix>/YYYY-MM-DD.csv
Re-running a day replaces that day's object. Finout ingests newly created or
updated files and replaces previously ingested data when an older file changes.
This makes retries and P7D backfills idempotent.
Rows whose LLM price cannot be resolved are skipped with a warning instead of being reported as zero cost. The affected day is marked pending and re-exported automatically once pricing resolves (see LLM cost database: freshness and timing). Non-USD cost fails the run because Finout's custom cost contract expects USD. A completed day with no rows is still uploaded as a header-only CSV so a retry can clear stale data from an earlier export.
To connect the bucket, grant Finout read access to the configured prefix and follow its Custom Cost Center onboarding process. Finout currently asks customers to contact support with the S3-hosted sample CSV.
Build the image:
docker build -t ghcr.io/openmeterio/finout-cost-reporter:0.1.0 .Install the CronJob with chart-managed secrets:
helm upgrade --install finout-cost-reporter \
deploy/helm/finout-cost-reporter \
--namespace finout --create-namespace \
--set secret.apiToken="$API_TOKEN" \
--set secret.s3AccessKeyId="$S3_ACCESS_KEY_ID" \
--set secret.s3SecretAccessKey="$S3_SECRET_ACCESS_KEY" \
--set config.s3.bucket="my-finout-cost-bucket"For AWS workload identity, omit the S3 keys and annotate the service account:
secret:
apiToken: kpat_REPLACE_ME
serviceAccount:
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/finout-reporter
automountServiceAccountToken: trueThe reporter's S3 identity needs s3:PutObject (write each day's CSV and the
state marker) and s3:GetObject (read the state marker to resume) on
arn:aws:s3:::<bucket>/<prefix>/*. The GetObject permission is required. A
write-only policy makes every run re-bootstrap and can leave gaps. (Finout's own
read access to the bucket is separate. See below.)
For production secret management, set secret.create=false and
secret.existingSecret to a Secret containing API_TOKEN. Static S3 keys are
optional.
The default schedule is 01:00 UTC. P1D always queries the previous complete
UTC day, regardless of the actual start time. concurrencyPolicy: Forbid
prevents overlapping exports.
pnpm check # format check + lint + typecheck + test (the CI gate)
pnpm fmt # auto-format with oxfmt
pnpm lint:fix # auto-fix lint issues with oxlint
pnpm typecheck
pnpm test
helm lint deploy/helm/finout-cost-reporter \
-f deploy/helm/finout-cost-reporter/ci/test-values.yaml
helm template test deploy/helm/finout-cost-reporter \
-f deploy/helm/finout-cost-reporter/ci/test-values.yamlCode is formatted with oxfmt
and linted with oxlint. One enforced rule: if/else/loop
bodies are always braced (curly).
API shapes are provided by @openmeter/client, generated from the
Konnect Metering & Billing OpenAPI specification.
Garage configuration follows the
devenv Garage service.