12 KiB
| name | description | category |
|---|---|---|
| redis-atomic-balance | Redis atomic balance reservation for concurrent LLM request deduplication | devops |
Problem
checkCustomerBalance reads DB balance, then consume_accounting deducts later in
background loop. Concurrent requests all pass the check before any deduction runs,
enabling overspend. At 2000 concurrency, this race window is dangerously wide.
Solution
Three-phase Redis atomic flow integrated with product_management accounting:
Phase 1 — RESERVE (before inference)
Redis Lua: GET model:max_cost:{llmid}
(if missing, read max from DB llmusage history)
balance = GET balance:{userorgid}
(if missing, read from accounting.getCustomerBalance DB)
if balance < max_cost → 429
DECRBY balance:{userorgid} max_cost
SETEX reserve:{luid} 600 {userorgid}|{llmid}|{max_cost}
Phase 2 — FINALIZE (after product_accounting/consume_accounting)
INCRBY balance:{userorgid} (max_cost - actual_cost)
if actual_cost > max_cost:
HSET model:max_cost {llmid} {actual_cost}
UPDATE llm SET max_cost={actual_cost} WHERE id={llmid}
DEL reserve:{luid}
Phase 3 — REFUND (on API failure)
INCRBY balance:{userorgid} max_cost
DEL reserve:{luid}
Integration points (as built 2026-08, deployed to test server)
| When | Where | What |
|---|---|---|
| Request arrival | llmage/llmclient.py _inference_generator |
Single unified reserve injection covering stream/sync/async and ALL 22 dspy entries. If params_kw._luid already set (chat/completions pre-reserved), skip and reuse — no double deduction |
| luid identity | all request functions | luid = params_kw.get('_luid') or getID() → reserve key == llmusage.id == finalize key (fixes the original luid mismatch) |
| Stream failure | chat/completions dspy except block + llmclient stream except | env.refund_balance(luid) |
| Sync failure | syncinference.py outer except |
if luid: await refund_balance(...) |
| Async failure | asyncinference.py |
refund on uapi.call submit failure, outer exception, and query_task_status FAILED (asynctask_callbacka hook does not exist — exhaustive search confirmed) |
| Product path | product_interface.py |
reserve + refund on failure, both stream and non-stream product execution |
| Accounting done | product_management/core.py:product_accounting AND llmage/accounting.py:llm_accounting |
finalize_balance after charging succeeds; pass llmid so cold-start max_cost updates even without a reserve |
| Recharge | external accounting repo recharge.py RechargeBiz.accounting |
invalidate_balance_cache() after successful charge — otherwise stale Redis balance keeps blocking newly-funded users |
Async tasks live longer than the 600s reserve TTL: extend_reserve prolongs it to 3600s. Accounting SQL only collects SUCCEEDED rows, so FAILED refunds never collide with charging.
Live runtime observations (verified 2026-08-05, test server 120.48.168.15:9180)
One smoke request (customer org, qwen3-max, llmid BDfl0Ts-sZp_UgRDGfDM4) produces exactly three keys:
balance:{userorgid} = 285442— integer cents, TTL=-1 (DB balance loaded minus one reserve)model:max_cost:{llmid} = 65— cents, TTL=-1reserve:{luid} = {userorgid}|{llmid}|65— TTL ≈588s (600 at creation)
backend_accounting loop sleeps 10s per iteration — finalize lands ~10-25s after a SUCCEEDED llmusage. In tests, wait ≥25s before asserting reserve cleanup. Redis URL comes from config.website.session_redis.url (redis://127.0.0.1:6379 on test). After a deploy/restart the balance/max_cost keys are legitimately zero until the first priced request — they have no TTL and don't vanish on their own.
Data-model facts for picking test models
llmtable has noppidand nostreamcolumn (cols: id/name/model/description/iconid/upappid/providerid/ownerid/enabled_date/expired_date/min_balance/status)ppidlives inllm_api_map(llmid, apiname, ppid, query_apiname, isdefaultcatelog)- stream mode lives in
uapi.stream('stream'|'sync'|'async'|'false'|NULL), joined viallm_api_map.apiname = uapi.name;get_llm()merges it at runtime (llmage/utils.pyllm.stream = uapi.stream) /llmage/v1/models/index.dspyreturns models with emptymode— stream mode cannot be read from the API; query the DBusers.orgidis the balance userorgid; platform adminswordhas orgid=0 (no org, no acc_balance row) — balance tests must use a real customer account (user3)llmusage.accounting_statusgoes created → accounted; only SUCCEEDED rows get chargedllmcatelogis just a directory table (id/name/description/hfid/ioid) — not where pricing lives
Audit / concurrency-test workflow (EXECUTED 2026-08-05, tokentest.opencomputing.cn)
Goal: prove atomic pre-deduct under concurrency and zero overdraft. Full phased script: references/balance-concurrency-audit-script.md — deployed on the server as /tmp/sage_loadtest.py, run with the venv python: cd /tmp && /d/apitest/sage/py3/bin/python sage_loadtest.py A|B|C1|C3|D. Executed 2026-08-05 after user approval ("可以在tokentest.opencomputing.cn上测试"); the audit caught two real blockers: the GETDEL/Redis-6.0 bug (Pitfall 3) and sqlor MVCC snapshot pinning blinding backend_accounting (Pitfall 6).
- A full-balance ✅ PASS: 1 warmup + N=20 concurrent qwen3-max, sampler polls Redis. Result: 20/20 ok, min_balance = base − N×max_cost (never negative over 55 samples), peak reserves = N+1 (warmup's included), reserves all drained after ~25s finalize wait, balance restored — reconcile delta matched actual costs.
- B overdraft guard ✅ PASS (behaviorally):
SET balance:{org} 200(max_cost=65 → floor 3 fit) → 10 concurrent → exactly 3 ok, 7 rejected. Rejection shape is HTTP 429 with JSON body"code": "insufficient_quota"("You exceeded your current quota") — NOT the余额不足substring; detect on 429/insufficient_quota. min_balance=5 (200−3×65), never negative; finalize restores balance to exactly 200; DEL the key afterwards to restore DB-reload behavior. - C1 FAILED refund ❌ not triggered:
max_tokens=99999999was ACCEPTED upstream (HTTP 200, normal completion) — this injection vector does not work with the current provider. Need a different failure vector (bad model name, dead sync endpoint, or upstream-side error). - C3 invalidate and D consistency: script ready, run as listed.
Login/cookie facts (corrected from earlier notes): login endpoint is POST /rbac/user/up_login.dspy with JSON {"username": ..., "password": ...} (plaintext password — the dspy does RC4 internally, do not pre-encode), cookie jar at /tmp/u3_cookie.txt. Sessions expire mid-test (all calls return 401) — re-login before each phase. Chat path: POST /llmage/v1/chat/completions/index.dspy.
Workflow rule: these tests spend REAL customer balance (user3) and hit REAL upstream LLM providers. Get explicit user approval before running them. The user denied the upload once (2026-08-05 morning) and approved later the same day ("可以在tokentest.opencomputing.cn上测试") — on denial, stop and ask; never rephrase or reroute around it.
Crash recovery
Redis restart → wait for backend_accounting loop to complete all in-flight
→ rebuild balance:{userorgid} from DB accounting records
→ rebuild model:max_cost:{llmid} from llm.max_cost column
→ clear all reserve:* keys (TTL 10min auto-clears if missed)
Guard matrix
| Scenario | Protection |
|---|---|
| Double finalize | DEL reserve:{luid} is atomic — second call returns null, skip |
| Reserve after crash | TTL=600s auto-rollback; on restart, clear all reserve:* |
| Redis down | balance.py now owns its connection singleton (no more silent no-op from missing env.redis). Verify on deploy that ConnectionError/TimeoutError map to no_redis so the DSPY DB-fallback fires instead of a false 429 (original Pitfall #1) |
| max_cost drifts up | finalize updates Redis AND DB when actual > max |
Pitfalls (from 2026-07 critical review)
- RESOLVED 2026-08:
env.rediswas never injected. The original design assumed externalenv.redisinjection — exhaustive search proved no code ever assigns it, which silently made every reserve a no-op (the root cause of concurrent overspend). Fix shipped:balance.pyowns a module-levelredis.asynciosingleton (URL fromconfig.jsonwebsite.session_redis.url, connection pattern copied fromappPublic/share_cache.py). The tpac-user / self-owned-org skip policy is centralized insidereserve_balance(useridparameter). Pyright's 3 errors on the redis calls are known type-stub false positives withdecode_responses=True— same pattern runs in production via share_cache. The oldno_redisDB-fallback branch in the DSPY is now a last-resort path only. - Reserve over-deduct race (permanent overcharge). RESERVE_LUA deducts
math.max(max_cost, stored_max)but thereserve:{luid}value stores the Python-side max_cost. If finalize updatesmodel:max_cost:{llmid}between Python's GET and the Lua EVAL, Lua deducts more than the stored reserve records; finalize then refunds against the lower stored value, so the excess is never returned (until balance key re-seeds from DB). Fix: buildreserve_valinside Lua using the final adjusted max_cost, not in Python. - GETDEL needs Redis ≥ 6.2 — RESOLVED 2026-08-05. The test server runs Redis 6.0.16, so GETDEL raised "Unknown Redis command" inside FINALIZE_LUA/REFUND_LUA: reserves were never released, balance never recovered, and all 20 Phase-A requests leaked their reserve. Fix (llmage commit 31c8fde): replace GETDEL with GET+DEL in both Lua scripts. Lesson: verify Redis version at deploy (
redis-server --version) and never use ≥6.2-only commands in Lua without a fallback. - Leaks are bounded by TTL only. Exceptions between reserve and generator start (e.g.
checkCustomerBalanceraising in the DSPY) never refund — the 600s TTL is the only cleanup. Acceptable, but don't extend TTL without adding a refund path. _centsuses banker's rounding (roundis round-half-even): exact half-cent values under-charge by 0.01. Rare; note if amounts can land on 0.005 boundaries.- Accounting-loop blindness silently breaks finalize (sqlor MVCC snapshot pinning). RESOLVED 2026-08-05. If
backend_accountinglogs "got 0 records" forever while llmusage rows exist, finalize never runs → reserves leak by TTL only and balance never recovers — the symptom looks like a reserve bug but the cause is the loop's DB connection holding a never-committed read transaction (REPEATABLE READ snapshot pinned at process start). Diagnosis:SELECT trx_state, trx_started, TIMESTAMPDIFF(SECOND,trx_started,NOW()), trx_mysql_thread_id, trx_query FROM information_schema.INNODB_TRX— a RUNNING trx with age ≈ process uptime andtrx_query=NULLis the smoking gun. Fix: sqlormysqlor.enter()commits before creating the cursor (commit fab420c — was already in the local repo for weeks but undeployed on the test server; check server rev before re-diagnosing). Full recipe: see theasync-db-connection-pool-reliabilityskill,references/mvcc-snapshot-pinning.md. Also: a stalebalance:{org}key poisoned by an old over-deduct bug must be DELeted once so it re-seeds from DB — it never self-heals (no TTL).
Files (as built 2026-08)
llmage/llmage/balance.py— reserve/finalize/refund + Lua scripts,extend_reserve,invalidate_balance_cache, module-level redis.asyncio singletonllmage/llmage/llmclient.py— unified reserve injection in_inference_generator, luid reuse, stream-exception refundllmage/llmage/syncinference.py/asyncinference.py— luid reuse + failure refundsllmage/llmage/product_interface.py— product-path reserve/refundllmage/llmage/accounting.py— finalize afterllm_accounting(direct-llmage charging path)llmage/llmage/init.py— env lambda wiring (incl. ttl param)llmage/wwwroot/v1/chat/completions/index.dspy— reserve call passesuseridproduct_management/core.py— finalize passesllmid- external
accounting/recharge.py— invalidate Redis balance cache after recharge
Deployment: Python-package changes need venv pip install + full restart (4 sage workers AND backend_accounting — see sage-module-deployment, incl. the bare-pip-installs-to-~/.local trap). The dspy change goes live on git pull alone (wwwroot symlink). Verify by grepping the venv site-packages for the new markers, not by pip output.