From 1abaa4ea2ece78f400b6a5d6adacbc336e3544cd Mon Sep 17 00:00:00 2001 From: yumoqing Date: Sun, 16 Aug 2026 15:17:43 +0800 Subject: [PATCH] =?UTF-8?q?feat:=20=E6=8A=80=E8=83=BD=E5=BA=93(skills=5Fli?= =?UTF-8?q?brary=20213=E4=B8=AAHermes=E6=8A=80=E8=83=BD)+skill=5Fpack?= =?UTF-8?q?=E6=8A=80=E8=83=BD=E9=9B=86=E5=AE=89=E8=A3=85=E8=83=BD=E5=8A=9B?= =?UTF-8?q?+ocai-h5-dev=E6=8A=80=E8=83=BD=E9=9B=86(37=E4=B8=AA=E5=BA=94?= =?UTF-8?q?=E7=94=A8/=E6=A8=A1=E5=9D=97=E5=BC=80=E5=8F=91=E6=8A=80?= =?UTF-8?q?=E8=83=BD)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- pipeline_core/skill_pack.py | 125 + .../all/accounting-module-example/SKILL.md | 190 ++ .../all/agentic-report-generation/SKILL.md | 133 + .../all/ahserver-hot-reload/SKILL.md | 449 ++++ skills_library/all/ahserver-pitfalls/SKILL.md | 318 +++ .../all/ahserver-post-debugging/SKILL.md | 51 + skills_library/all/ahserver/SKILL.md | 1295 +++++++++ skills_library/all/ai-coding-agents/SKILL.md | 103 + .../all/ai-music-production/SKILL.md | 107 + skills_library/all/airtable/SKILL.md | 229 ++ skills_library/all/api-load-testing/SKILL.md | 1158 ++++++++ .../all/appbase-module-example/SKILL.md | 145 + .../all/apppublic-python-module/SKILL.md | 159 ++ .../all/architecture-diagram/SKILL.md | 148 + skills_library/all/arxiv/SKILL.md | 282 ++ skills_library/all/ascii-art/SKILL.md | 322 +++ skills_library/all/ascii-video/SKILL.md | 241 ++ .../SKILL.md | 234 ++ skills_library/all/audit-log-module/SKILL.md | 156 ++ skills_library/all/audit-logging/SKILL.md | 97 + skills_library/all/auto-model-config/SKILL.md | 494 ++++ .../SKILL.md | 142 + skills_library/all/axolotl/SKILL.md | 165 ++ .../all/baoyu-article-illustrator/SKILL.md | 207 ++ skills_library/all/baoyu-comic/SKILL.md | 247 ++ skills_library/all/baoyu-infographic/SKILL.md | 237 ++ skills_library/all/bidding-documents/SKILL.md | 184 ++ skills_library/all/blogwatcher/SKILL.md | 137 + .../all/bricks-app-checklist/SKILL.md | 193 ++ .../all/bricks-chart-widgets/SKILL.md | 131 + .../all/bricks-dev-patterns/SKILL.md | 157 ++ skills_library/all/bricks-framework/SKILL.md | 1857 +++++++++++++ .../all/bricks-layout-patterns/SKILL.md | 310 +++ .../all/bricks-menu-dspy-pitfalls/SKILL.md | 63 + .../all/bricks-terminal-and-popup/SKILL.md | 57 + skills_library/all/bricks-ui-testing/SKILL.md | 213 ++ .../all/bricks-widget-development/SKILL.md | 1064 ++++++++ .../all/bricks-wterm-terminal/SKILL.md | 130 + .../all/browser-app-testing/SKILL.md | 298 +++ .../all/browser-automation/SKILL.md | 316 +++ skills_library/all/browser-harness/SKILL.md | 49 + skills_library/all/browser-setup/SKILL.md | 128 + .../all/build-script-modularization/SKILL.md | 151 ++ skills_library/all/claude-design/SKILL.md | 607 +++++ .../all/cli-process-loop-reliability/SKILL.md | 61 + skills_library/all/clip/SKILL.md | 256 ++ skills_library/all/cockpit-agent-dev/SKILL.md | 159 ++ .../all/cockpit-agent-patterns/SKILL.md | 216 ++ .../all/codebase-inspection/SKILL.md | 116 + skills_library/all/comfyui/SKILL.md | 612 +++++ skills_library/all/computer-use/SKILL.md | 356 +++ skills_library/all/creative-ideation/SKILL.md | 152 ++ .../all/crud-definition-spec/SKILL.md | 1506 +++++++++++ .../all/customer-facing-product-docs/SKILL.md | 39 + .../database-table-definition-spec/SKILL.md | 381 +++ skills_library/all/design-md/SKILL.md | 220 ++ skills_library/all/docx/SKILL.md | 127 + skills_library/all/dogfood/SKILL.md | 162 ++ .../dspy-file-implementation-spec/SKILL.md | 973 +++++++ skills_library/all/dspy-patterns/SKILL.md | 84 + skills_library/all/dspy/SKILL.md | 594 ++++ .../all/dynamic-page-extract/SKILL.md | 170 ++ .../all/email-server-setup/SKILL.md | 201 ++ .../all/evaluating-llms-harness/SKILL.md | 498 ++++ skills_library/all/excalidraw/SKILL.md | 199 ++ skills_library/all/find-nearby/SKILL.md | 69 + .../all/freelance-platforms/SKILL.md | 94 + skills_library/all/gguf/SKILL.md | 430 +++ skills_library/all/gif-search/SKILL.md | 91 + skills_library/all/git-mirroring/SKILL.md | 129 + skills_library/all/github-workflow/SKILL.md | 87 + skills_library/all/godmode/SKILL.md | 404 +++ skills_library/all/google-workspace/SKILL.md | 335 +++ .../all/gpu-async-service-pattern/SKILL.md | 313 +++ .../all/gpu-server-services/SKILL.md | 203 ++ .../all/grounded-citations/SKILL.md | 232 ++ skills_library/all/grpo-rl-training/SKILL.md | 575 ++++ skills_library/all/guidance/SKILL.md | 575 ++++ .../SKILL.md | 219 ++ .../all/harnessed-module-development/SKILL.md | 2269 ++++++++++++++++ skills_library/all/hermes-agent/SKILL.md | 1060 ++++++++ skills_library/all/hermes-app-deploy/SKILL.md | 474 ++++ .../all/hermes-cli-maintenance/SKILL.md | 107 + .../SKILL.md | 789 ++++++ .../all/hermes-state-backup/SKILL.md | 131 + .../all/hermes-state-merge/SKILL.md | 221 ++ .../hermes-web-cli-main-architecture/SKILL.md | 203 ++ .../SKILL.md | 201 ++ skills_library/all/himalaya/SKILL.md | 304 +++ skills_library/all/huggingface-hub/SKILL.md | 81 + skills_library/all/humanizer/SKILL.md | 647 +++++ .../inspecting-hermes-desktop-dom/SKILL.md | 159 ++ .../all/integrated-crm-app/SKILL.md | 568 ++++ .../all/json-to-ddl-generator/SKILL.md | 181 ++ .../all/jupyter-live-kernel/SKILL.md | 167 ++ .../all/kanban-orchestrator/SKILL.md | 284 ++ skills_library/all/kanban-worker/SKILL.md | 192 ++ .../SKILL.md | 540 ++++ .../all/ktv-video-production/SKILL.md | 1231 +++++++++ skills_library/all/linear/SKILL.md | 380 +++ skills_library/all/llama-cpp/SKILL.md | 249 ++ .../all/llm-api-config-from-url/SKILL.md | 538 ++++ skills_library/all/llm-wiki/SKILL.md | 507 ++++ .../all/llmage-api-testing/SKILL.md | 665 +++++ skills_library/all/llmage-module/SKILL.md | 1036 +++++++ skills_library/all/lyric-evaluator/SKILL.md | 99 + .../all/mail-server-deployment/SKILL.md | 214 ++ skills_library/all/mail-server-setup/SKILL.md | 253 ++ skills_library/all/manim-video/SKILL.md | 269 ++ skills_library/all/maps/SKILL.md | 195 ++ skills_library/all/mcporter/SKILL.md | 122 + .../all/minecraft-modpack-server/SKILL.md | 187 ++ skills_library/all/modal/SKILL.md | 344 +++ .../all/model-cost-optimization/SKILL.md | 80 + .../all/module-development-spec/SKILL.md | 1699 ++++++++++++ .../all/module-git-sync-workflow/SKILL.md | 186 ++ .../all/multi-agent-workflow/SKILL.md | 132 + .../all/multimodal-ai-inference-spec/SKILL.md | 357 +++ skills_library/all/native-mcp/SKILL.md | 357 +++ .../all/no-sudo-server-setup/SKILL.md | 108 + skills_library/all/notion/SKILL.md | 448 ++++ skills_library/all/obliteratus/SKILL.md | 342 +++ skills_library/all/obsidian/SKILL.md | 61 + skills_library/all/ocr-and-documents/SKILL.md | 196 ++ skills_library/all/officecli/SKILL.md | 96 + skills_library/all/openhue/SKILL.md | 112 + skills_library/all/outlines/SKILL.md | 655 +++++ skills_library/all/p5js/SKILL.md | 556 ++++ .../all/pccs-deploy-patterns/SKILL.md | 250 ++ skills_library/all/pccs-deploy/SKILL.md | 310 +++ skills_library/all/pccs-ui-design/SKILL.md | 125 + skills_library/all/pdf/SKILL.md | 174 ++ skills_library/all/peft/SKILL.md | 434 +++ skills_library/all/petdex/SKILL.md | 89 + skills_library/all/pipeline-agent-v2/SKILL.md | 553 ++++ .../all/pipeline-app-module/SKILL.md | 915 +++++++ .../all/pipeline-sage-bridge/SKILL.md | 62 + .../all/pipeline-sdlc-workflow/SKILL.md | 246 ++ skills_library/all/pixel-art/SKILL.md | 218 ++ skills_library/all/pokemon-player/SKILL.md | 216 ++ skills_library/all/polymarket/SKILL.md | 77 + .../all/popular-web-designs/SKILL.md | 214 ++ skills_library/all/powerpoint/SKILL.md | 256 ++ .../all/pre-commit-crud-check/SKILL.md | 110 + skills_library/all/pretext/SKILL.md | 220 ++ .../all/pricing-data-format/SKILL.md | 709 +++++ skills_library/all/pricing-module/SKILL.md | 875 ++++++ .../all/production-llm-serving/SKILL.md | 221 ++ .../python-module-import-debugging/SKILL.md | 192 ++ skills_library/all/pytorch-fsdp/SKILL.md | 129 + .../all/rag-new-dspy-permission/SKILL.md | 28 + skills_library/all/rag-operations/SKILL.md | 183 ++ .../all/ragserver-development/SKILL.md | 1762 ++++++++++++ .../SKILL.md | 925 +++++++ .../all/reallife_asset-downapp-api/SKILL.md | 362 +++ .../all/redis-atomic-balance/SKILL.md | 142 + .../all/reference-module-protection/SKILL.md | 208 ++ .../all/remote-app-testing/SKILL.md | 340 +++ .../all/requesting-code-review/SKILL.md | 280 ++ .../all/research-paper-writing/SKILL.md | 2377 +++++++++++++++++ .../all/sage-app-deploy-checklist/SKILL.md | 197 ++ .../all/sage-bricks-app-patterns/SKILL.md | 178 ++ .../all/sage-dspy-development/SKILL.md | 497 ++++ skills_library/all/sage-frontend/SKILL.md | 243 ++ .../all/sage-module-deployment/SKILL.md | 684 +++++ .../all/sage-module-scaffolding/SKILL.md | 159 ++ .../all/sage-performance-profiling/SKILL.md | 180 ++ skills_library/all/sage-platform/SKILL.md | 808 ++++++ .../all/sandboxed-deployment/SKILL.md | 118 + .../all/sdlc-agent-collaboration/SKILL.md | 519 ++++ .../all/sdlc-repo-standard/SKILL.md | 101 + skills_library/all/segment-anything/SKILL.md | 506 ++++ skills_library/all/serving-llms-vllm/SKILL.md | 373 +++ skills_library/all/showcase-module/SKILL.md | 87 + skills_library/all/simplify-code/SKILL.md | 270 ++ .../all/socks-proxy-download/SKILL.md | 170 ++ .../SKILL.md | 313 +++ skills_library/all/spike/SKILL.md | 197 ++ skills_library/all/spotify/SKILL.md | 135 + .../SKILL.md | 215 ++ .../all/sqlor-database-module/SKILL.md | 1072 ++++++++ skills_library/all/ssh-password-auth/SKILL.md | 88 + skills_library/all/stable-diffusion/SKILL.md | 522 ++++ .../all/standalone-sage-app-deploy/SKILL.md | 130 + .../all/standalone-sage-deployment/SKILL.md | 62 + .../student-grade-management-example/SKILL.md | 96 + .../all/suno-prompt-engineer/SKILL.md | 138 + .../all/supplychain-pitfalls/SKILL.md | 1314 +++++++++ .../all/systematic-debugging/SKILL.md | 608 +++++ .../all/task-and-input-logging/SKILL.md | 54 + skills_library/all/task-reliability/SKILL.md | 278 ++ .../all/test-driven-development/SKILL.md | 362 +++ skills_library/all/touchdesigner-mcp/SKILL.md | 356 +++ skills_library/all/trl-fine-tuning/SKILL.md | 462 ++++ .../all/uapi-integration-patterns/SKILL.md | 52 + skills_library/all/uapi-module/SKILL.md | 110 + .../all/ui-ux-pro-max-bricks/SKILL.md | 359 +++ skills_library/all/unipay/SKILL.md | 89 + skills_library/all/unsloth/SKILL.md | 83 + .../all/vendor-pricing-sql-generator/SKILL.md | 683 +++++ .../all/vllm-multi-instance-gpu/SKILL.md | 271 ++ skills_library/all/voucher-module/SKILL.md | 145 + .../all/web-application-spec/SKILL.md | 290 ++ skills_library/all/webapp-deploy/SKILL.md | 219 ++ .../all/webapp-remote-deploy/SKILL.md | 149 ++ .../all/weights-and-biases/SKILL.md | 598 +++++ skills_library/all/whisper/SKILL.md | 320 +++ skills_library/all/writing-plans/SKILL.md | 301 +++ skills_library/all/xitter/SKILL.md | 202 ++ skills_library/all/xlsx/SKILL.md | 105 + skills_library/all/xurl/SKILL.md | 436 +++ skills_library/all/youtube-content/SKILL.md | 76 + skills_library/all/yuanbao/SKILL.md | 108 + .../all/zero-root-bwrap-sandbox/SKILL.md | 162 ++ .../packs/ocai-h5-dev/manifest.json | 47 + 215 files changed, 74068 insertions(+) create mode 100644 pipeline_core/skill_pack.py create mode 100644 skills_library/all/accounting-module-example/SKILL.md create mode 100644 skills_library/all/agentic-report-generation/SKILL.md create mode 100644 skills_library/all/ahserver-hot-reload/SKILL.md create mode 100644 skills_library/all/ahserver-pitfalls/SKILL.md create mode 100644 skills_library/all/ahserver-post-debugging/SKILL.md create mode 100644 skills_library/all/ahserver/SKILL.md create mode 100644 skills_library/all/ai-coding-agents/SKILL.md create mode 100644 skills_library/all/ai-music-production/SKILL.md create mode 100644 skills_library/all/airtable/SKILL.md create mode 100644 skills_library/all/api-load-testing/SKILL.md create mode 100644 skills_library/all/appbase-module-example/SKILL.md create mode 100644 skills_library/all/apppublic-python-module/SKILL.md create mode 100644 skills_library/all/architecture-diagram/SKILL.md create mode 100644 skills_library/all/arxiv/SKILL.md create mode 100644 skills_library/all/ascii-art/SKILL.md create mode 100644 skills_library/all/ascii-video/SKILL.md create mode 100644 skills_library/all/async-db-connection-pool-reliability/SKILL.md create mode 100644 skills_library/all/audit-log-module/SKILL.md create mode 100644 skills_library/all/audit-logging/SKILL.md create mode 100644 skills_library/all/auto-model-config/SKILL.md create mode 100644 skills_library/all/automated-video-production-pipeline/SKILL.md create mode 100644 skills_library/all/axolotl/SKILL.md create mode 100644 skills_library/all/baoyu-article-illustrator/SKILL.md create mode 100644 skills_library/all/baoyu-comic/SKILL.md create mode 100644 skills_library/all/baoyu-infographic/SKILL.md create mode 100644 skills_library/all/bidding-documents/SKILL.md create mode 100644 skills_library/all/blogwatcher/SKILL.md create mode 100644 skills_library/all/bricks-app-checklist/SKILL.md create mode 100644 skills_library/all/bricks-chart-widgets/SKILL.md create mode 100644 skills_library/all/bricks-dev-patterns/SKILL.md create mode 100644 skills_library/all/bricks-framework/SKILL.md create mode 100644 skills_library/all/bricks-layout-patterns/SKILL.md create mode 100644 skills_library/all/bricks-menu-dspy-pitfalls/SKILL.md create mode 100644 skills_library/all/bricks-terminal-and-popup/SKILL.md create mode 100644 skills_library/all/bricks-ui-testing/SKILL.md create mode 100644 skills_library/all/bricks-widget-development/SKILL.md create mode 100644 skills_library/all/bricks-wterm-terminal/SKILL.md create mode 100644 skills_library/all/browser-app-testing/SKILL.md create mode 100644 skills_library/all/browser-automation/SKILL.md create mode 100644 skills_library/all/browser-harness/SKILL.md create mode 100644 skills_library/all/browser-setup/SKILL.md create mode 100644 skills_library/all/build-script-modularization/SKILL.md create mode 100644 skills_library/all/claude-design/SKILL.md create mode 100644 skills_library/all/cli-process-loop-reliability/SKILL.md create mode 100644 skills_library/all/clip/SKILL.md create mode 100644 skills_library/all/cockpit-agent-dev/SKILL.md create mode 100644 skills_library/all/cockpit-agent-patterns/SKILL.md create mode 100644 skills_library/all/codebase-inspection/SKILL.md create mode 100644 skills_library/all/comfyui/SKILL.md create mode 100644 skills_library/all/computer-use/SKILL.md create mode 100644 skills_library/all/creative-ideation/SKILL.md create mode 100644 skills_library/all/crud-definition-spec/SKILL.md create mode 100644 skills_library/all/customer-facing-product-docs/SKILL.md create mode 100644 skills_library/all/database-table-definition-spec/SKILL.md create mode 100644 skills_library/all/design-md/SKILL.md create mode 100644 skills_library/all/docx/SKILL.md create mode 100644 skills_library/all/dogfood/SKILL.md create mode 100644 skills_library/all/dspy-file-implementation-spec/SKILL.md create mode 100644 skills_library/all/dspy-patterns/SKILL.md create mode 100644 skills_library/all/dspy/SKILL.md create mode 100644 skills_library/all/dynamic-page-extract/SKILL.md create mode 100644 skills_library/all/email-server-setup/SKILL.md create mode 100644 skills_library/all/evaluating-llms-harness/SKILL.md create mode 100644 skills_library/all/excalidraw/SKILL.md create mode 100644 skills_library/all/find-nearby/SKILL.md create mode 100644 skills_library/all/freelance-platforms/SKILL.md create mode 100644 skills_library/all/gguf/SKILL.md create mode 100644 skills_library/all/gif-search/SKILL.md create mode 100644 skills_library/all/git-mirroring/SKILL.md create mode 100644 skills_library/all/github-workflow/SKILL.md create mode 100644 skills_library/all/godmode/SKILL.md create mode 100644 skills_library/all/google-workspace/SKILL.md create mode 100644 skills_library/all/gpu-async-service-pattern/SKILL.md create mode 100644 skills_library/all/gpu-server-services/SKILL.md create mode 100644 skills_library/all/grounded-citations/SKILL.md create mode 100644 skills_library/all/grpo-rl-training/SKILL.md create mode 100644 skills_library/all/guidance/SKILL.md create mode 100644 skills_library/all/harnessed-agent-skill-architecture/SKILL.md create mode 100644 skills_library/all/harnessed-module-development/SKILL.md create mode 100644 skills_library/all/hermes-agent/SKILL.md create mode 100644 skills_library/all/hermes-app-deploy/SKILL.md create mode 100644 skills_library/all/hermes-cli-maintenance/SKILL.md create mode 100644 skills_library/all/hermes-service-module-implementation/SKILL.md create mode 100644 skills_library/all/hermes-state-backup/SKILL.md create mode 100644 skills_library/all/hermes-state-merge/SKILL.md create mode 100644 skills_library/all/hermes-web-cli-main-architecture/SKILL.md create mode 100644 skills_library/all/hermes-web-ui-build-troubleshooting/SKILL.md create mode 100644 skills_library/all/himalaya/SKILL.md create mode 100644 skills_library/all/huggingface-hub/SKILL.md create mode 100644 skills_library/all/humanizer/SKILL.md create mode 100644 skills_library/all/inspecting-hermes-desktop-dom/SKILL.md create mode 100644 skills_library/all/integrated-crm-app/SKILL.md create mode 100644 skills_library/all/json-to-ddl-generator/SKILL.md create mode 100644 skills_library/all/jupyter-live-kernel/SKILL.md create mode 100644 skills_library/all/kanban-orchestrator/SKILL.md create mode 100644 skills_library/all/kanban-worker/SKILL.md create mode 100644 skills_library/all/kotlin-multiplatform-compose-desktop/SKILL.md create mode 100644 skills_library/all/ktv-video-production/SKILL.md create mode 100644 skills_library/all/linear/SKILL.md create mode 100644 skills_library/all/llama-cpp/SKILL.md create mode 100644 skills_library/all/llm-api-config-from-url/SKILL.md create mode 100644 skills_library/all/llm-wiki/SKILL.md create mode 100644 skills_library/all/llmage-api-testing/SKILL.md create mode 100644 skills_library/all/llmage-module/SKILL.md create mode 100644 skills_library/all/lyric-evaluator/SKILL.md create mode 100644 skills_library/all/mail-server-deployment/SKILL.md create mode 100644 skills_library/all/mail-server-setup/SKILL.md create mode 100644 skills_library/all/manim-video/SKILL.md create mode 100644 skills_library/all/maps/SKILL.md create mode 100644 skills_library/all/mcporter/SKILL.md create mode 100644 skills_library/all/minecraft-modpack-server/SKILL.md create mode 100644 skills_library/all/modal/SKILL.md create mode 100644 skills_library/all/model-cost-optimization/SKILL.md create mode 100644 skills_library/all/module-development-spec/SKILL.md create mode 100644 skills_library/all/module-git-sync-workflow/SKILL.md create mode 100644 skills_library/all/multi-agent-workflow/SKILL.md create mode 100644 skills_library/all/multimodal-ai-inference-spec/SKILL.md create mode 100644 skills_library/all/native-mcp/SKILL.md create mode 100644 skills_library/all/no-sudo-server-setup/SKILL.md create mode 100644 skills_library/all/notion/SKILL.md create mode 100644 skills_library/all/obliteratus/SKILL.md create mode 100644 skills_library/all/obsidian/SKILL.md create mode 100644 skills_library/all/ocr-and-documents/SKILL.md create mode 100644 skills_library/all/officecli/SKILL.md create mode 100644 skills_library/all/openhue/SKILL.md create mode 100644 skills_library/all/outlines/SKILL.md create mode 100644 skills_library/all/p5js/SKILL.md create mode 100644 skills_library/all/pccs-deploy-patterns/SKILL.md create mode 100644 skills_library/all/pccs-deploy/SKILL.md create mode 100644 skills_library/all/pccs-ui-design/SKILL.md create mode 100644 skills_library/all/pdf/SKILL.md create mode 100644 skills_library/all/peft/SKILL.md create mode 100644 skills_library/all/petdex/SKILL.md create mode 100644 skills_library/all/pipeline-agent-v2/SKILL.md create mode 100644 skills_library/all/pipeline-app-module/SKILL.md create mode 100644 skills_library/all/pipeline-sage-bridge/SKILL.md create mode 100644 skills_library/all/pipeline-sdlc-workflow/SKILL.md create mode 100644 skills_library/all/pixel-art/SKILL.md create mode 100644 skills_library/all/pokemon-player/SKILL.md create mode 100644 skills_library/all/polymarket/SKILL.md create mode 100644 skills_library/all/popular-web-designs/SKILL.md create mode 100644 skills_library/all/powerpoint/SKILL.md create mode 100644 skills_library/all/pre-commit-crud-check/SKILL.md create mode 100644 skills_library/all/pretext/SKILL.md create mode 100644 skills_library/all/pricing-data-format/SKILL.md create mode 100644 skills_library/all/pricing-module/SKILL.md create mode 100644 skills_library/all/production-llm-serving/SKILL.md create mode 100644 skills_library/all/python-module-import-debugging/SKILL.md create mode 100644 skills_library/all/pytorch-fsdp/SKILL.md create mode 100644 skills_library/all/rag-new-dspy-permission/SKILL.md create mode 100644 skills_library/all/rag-operations/SKILL.md create mode 100644 skills_library/all/ragserver-development/SKILL.md create mode 100644 skills_library/all/rbac-permission-initialization-pattern/SKILL.md create mode 100644 skills_library/all/reallife_asset-downapp-api/SKILL.md create mode 100644 skills_library/all/redis-atomic-balance/SKILL.md create mode 100644 skills_library/all/reference-module-protection/SKILL.md create mode 100644 skills_library/all/remote-app-testing/SKILL.md create mode 100644 skills_library/all/requesting-code-review/SKILL.md create mode 100644 skills_library/all/research-paper-writing/SKILL.md create mode 100644 skills_library/all/sage-app-deploy-checklist/SKILL.md create mode 100644 skills_library/all/sage-bricks-app-patterns/SKILL.md create mode 100644 skills_library/all/sage-dspy-development/SKILL.md create mode 100644 skills_library/all/sage-frontend/SKILL.md create mode 100644 skills_library/all/sage-module-deployment/SKILL.md create mode 100644 skills_library/all/sage-module-scaffolding/SKILL.md create mode 100644 skills_library/all/sage-performance-profiling/SKILL.md create mode 100644 skills_library/all/sage-platform/SKILL.md create mode 100644 skills_library/all/sandboxed-deployment/SKILL.md create mode 100644 skills_library/all/sdlc-agent-collaboration/SKILL.md create mode 100644 skills_library/all/sdlc-repo-standard/SKILL.md create mode 100644 skills_library/all/segment-anything/SKILL.md create mode 100644 skills_library/all/serving-llms-vllm/SKILL.md create mode 100644 skills_library/all/showcase-module/SKILL.md create mode 100644 skills_library/all/simplify-code/SKILL.md create mode 100644 skills_library/all/socks-proxy-download/SKILL.md create mode 100644 skills_library/all/software-development__pipeline-agent-architecture/SKILL.md create mode 100644 skills_library/all/spike/SKILL.md create mode 100644 skills_library/all/spotify/SKILL.md create mode 100644 skills_library/all/sqlor-database-framework-enhancement/SKILL.md create mode 100644 skills_library/all/sqlor-database-module/SKILL.md create mode 100644 skills_library/all/ssh-password-auth/SKILL.md create mode 100644 skills_library/all/stable-diffusion/SKILL.md create mode 100644 skills_library/all/standalone-sage-app-deploy/SKILL.md create mode 100644 skills_library/all/standalone-sage-deployment/SKILL.md create mode 100644 skills_library/all/student-grade-management-example/SKILL.md create mode 100644 skills_library/all/suno-prompt-engineer/SKILL.md create mode 100644 skills_library/all/supplychain-pitfalls/SKILL.md create mode 100644 skills_library/all/systematic-debugging/SKILL.md create mode 100644 skills_library/all/task-and-input-logging/SKILL.md create mode 100644 skills_library/all/task-reliability/SKILL.md create mode 100644 skills_library/all/test-driven-development/SKILL.md create mode 100644 skills_library/all/touchdesigner-mcp/SKILL.md create mode 100644 skills_library/all/trl-fine-tuning/SKILL.md create mode 100644 skills_library/all/uapi-integration-patterns/SKILL.md create mode 100644 skills_library/all/uapi-module/SKILL.md create mode 100644 skills_library/all/ui-ux-pro-max-bricks/SKILL.md create mode 100644 skills_library/all/unipay/SKILL.md create mode 100644 skills_library/all/unsloth/SKILL.md create mode 100644 skills_library/all/vendor-pricing-sql-generator/SKILL.md create mode 100644 skills_library/all/vllm-multi-instance-gpu/SKILL.md create mode 100644 skills_library/all/voucher-module/SKILL.md create mode 100644 skills_library/all/web-application-spec/SKILL.md create mode 100644 skills_library/all/webapp-deploy/SKILL.md create mode 100644 skills_library/all/webapp-remote-deploy/SKILL.md create mode 100644 skills_library/all/weights-and-biases/SKILL.md create mode 100644 skills_library/all/whisper/SKILL.md create mode 100644 skills_library/all/writing-plans/SKILL.md create mode 100644 skills_library/all/xitter/SKILL.md create mode 100644 skills_library/all/xlsx/SKILL.md create mode 100644 skills_library/all/xurl/SKILL.md create mode 100644 skills_library/all/youtube-content/SKILL.md create mode 100644 skills_library/all/yuanbao/SKILL.md create mode 100644 skills_library/all/zero-root-bwrap-sandbox/SKILL.md create mode 100644 skills_library/packs/ocai-h5-dev/manifest.json diff --git a/pipeline_core/skill_pack.py b/pipeline_core/skill_pack.py new file mode 100644 index 0000000..2b48d60 --- /dev/null +++ b/pipeline_core/skill_pack.py @@ -0,0 +1,125 @@ +"""技能集(skill pack)管理 + 安装能力。 + +技能库结构(pipeline-core 仓库内): + skills_library/ + ├── all/ # 完整技能库(每个技能一个目录,含 SKILL.md) + │ ├── module-development-spec/SKILL.md + │ └── ... + └── packs/ # 技能集定义 + └── ocai-h5-dev/ + └── manifest.json # 元数据 + 引用的技能列表 + +安装 = 把技能集里引用的技能从 all/ 复制到目标技能根目录的 orgs/{org_id}/ 下, +skill_loader 的 org scope 会自动加载。 +""" +import os +import json +import shutil + + +def get_library_dir(): + """技能库根目录。优先环境变量 PIPELINE_SKILLS_LIBRARY,否则默认 pipeline-core 仓库内 skills_library/。""" + env = os.environ.get("PIPELINE_SKILLS_LIBRARY", "") + if env: + return os.path.normpath(env) + return os.path.normpath( + os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "skills_library") + ) + + +LIBRARY_DIR = get_library_dir() +PACKS_DIR = os.path.join(LIBRARY_DIR, "packs") +ALL_DIR = os.path.join(LIBRARY_DIR, "all") + + +def list_packs(): + """列出所有可安装的技能集(manifest 摘要)。""" + packs = [] + if not os.path.isdir(PACKS_DIR): + return packs + for d in sorted(os.listdir(PACKS_DIR)): + mf = os.path.join(PACKS_DIR, d, "manifest.json") + if os.path.isfile(mf): + try: + with open(mf, "r", encoding="utf-8") as f: + m = json.load(f) + packs.append({ + "name": m.get("name", d), + "title": m.get("title", d), + "description": m.get("description", ""), + "vendor": m.get("vendor", ""), + "version": m.get("version", "1.0.0"), + "skill_count": len(m.get("skills", [])), + }) + except Exception: + continue + return packs + + +def get_pack(pack_name): + """返回技能集 manifest(含技能列表),不存在返回 None。""" + mf = os.path.join(PACKS_DIR, pack_name, "manifest.json") + if not os.path.isfile(mf): + return None + with open(mf, "r", encoding="utf-8") as f: + return json.load(f) + + +def install_pack(pack_name, target_base_dir, org_id): + """把技能集安装到目标技能根目录的 orgs/{org_id}/ 下。 + + target_base_dir: skill_loader 的 base_dir(运行时技能根目录,如 .../skills) + org_id: 机构 ID(安装到该机构的 org scope) + 返回 {success, installed:[...], skipped:[...], error} + """ + pack = get_pack(pack_name) + if not pack: + return {"success": False, "error": f"技能集不存在: {pack_name}"} + + org_dir = os.path.join(target_base_dir, "orgs", str(org_id)) + os.makedirs(org_dir, exist_ok=True) + + installed, skipped = [], [] + for skill_name in pack.get("skills", []): + src = os.path.join(ALL_DIR, skill_name) + if not os.path.isdir(src): + skipped.append(skill_name) + continue + dst = os.path.join(org_dir, skill_name) + if os.path.exists(dst): + shutil.rmtree(dst) + shutil.copytree(src, dst) + installed.append(skill_name) + + return {"success": True, "installed": installed, "skipped": skipped} + + +def uninstall_pack(pack_name, target_base_dir, org_id): + """卸载技能集:删除该机构下该技能集引用的技能目录。""" + pack = get_pack(pack_name) + if not pack: + return {"success": False, "error": f"技能集不存在: {pack_name}"} + org_dir = os.path.join(target_base_dir, "orgs", str(org_id)) + removed = [] + for skill_name in pack.get("skills", []): + dst = os.path.join(org_dir, skill_name) + if os.path.isdir(dst): + shutil.rmtree(dst) + removed.append(skill_name) + return {"success": True, "removed": removed} + + +def installed_packs(target_base_dir, org_id): + """返回某机构已安装的技能集名列表。""" + packs = [] + for p in list_packs(): + pack = get_pack(p["name"]) + org_dir = os.path.join(target_base_dir, "orgs", str(org_id)) + cnt = 0 + for skill_name in pack.get("skills", []): + if os.path.isdir(os.path.join(org_dir, skill_name)): + cnt += 1 + if cnt > 0: + packs.append({"name": p["name"], "title": p["title"], + "installed_skills": cnt, "total_skills": p["skill_count"]}) + return packs diff --git a/skills_library/all/accounting-module-example/SKILL.md b/skills_library/all/accounting-module-example/SKILL.md new file mode 100644 index 0000000..23e53fd --- /dev/null +++ b/skills_library/all/accounting-module-example/SKILL.md @@ -0,0 +1,190 @@ +--- +name: accounting-module-example +version: 1.0.0 +description: Complete example of a compliant accounting module following the module development specification, demonstrating proper structure, initialization, CRUD definitions, and integration patterns. +trigger_conditions: + - User wants to understand how to implement a real-world module following the module-development-spec + - Need reference implementation for accounting/billing functionality + - Looking for examples of ServerEnv exposure, CRUD configuration, and module organization +--- + +# Accounting Module Example + +## Overview +This skill provides a complete reference implementation of an accounting module that fully complies with the module development specification. The accounting module demonstrates proper organization, initialization patterns, CRUD definitions, and integration with the ahserver ecosystem. + +## Module Structure Analysis + +### Core Directory Structure +``` +accounting/ # Main module directory +├── accounting/ # Python package +│ ├── __init__.py # Python package marker +│ ├── init.py # Module initialization (load_accounting function) +│ ├── *.py # Core business logic files +├── json/ # CRUD definition files (.json) +│ ├── account.json +│ ├── accounting_log.json +│ ├── subject.json +│ ├── acc_detail.json +│ ├── acc_balance.json +│ ├── accounting_config.json +│ └── account_config.json +├── models/ # Database table definitions (.xlsx format in this example) +│ ├── account.xlsx +│ ├── acc_balance.xlsx +│ ├── acc_detail.xlsx +│ ├── subject.xlsx +│ └── ... (other table definitions) +├── wwwroot/ # Frontend scripts and resources +│ ├── *.ui # Jinja2 template files +│ ├── *.dspy # Controlled Python scripts +│ └── imgs/ # Image assets +├── init/ # Initialization data (not present in this example) +├── setup.py # Python packaging (legacy format) +├── requirements.txt # Dependencies +└── README.md # Module documentation +``` + +## Key Implementation Patterns + +### 1. Module Initialization (init.py) +The `load_accounting()` function properly exposes all necessary components through ServerEnv: + +```python +def load_accounting(): + g = ServerEnv() + g.Accounting = Accounting # Configuration class + g.RechargeBiz = RechargeBiz # Business logic class + g.consume_accounting = consume_accounting # Async functions + g.write_bill = write_bill + g.openOwnerAccounts = openOwnerAccounts # Account opening functions + g.openProviderAccounts = openProviderAccounts + g.openResellerAccounts = openResellerAccounts + g.openCustomerAccounts = openCustomerAccounts + g.getAccountBalance = getAccountBalance # Balance query functions + g.getCustomerBalance = getCustomerBalance + g.getAccountByName = getAccountByName + g.get_account_total_amount = get_account_total_amount + g.recharge_accounting = recharge_accounting + g.get_accdetail = get_accdetail # Detail query functions + g.all_my_accounts = all_my_accounts + g.openRetailRelationshipAccounts = openRetailRelationshipAccounts +``` + +### 2. CRUD Definition Example (account.json) +Demonstrates list view configuration with subtables for related data: + +```json +{ + "tblname": "account", + "title": "科目", + "params": { + "sortby": "name", + "browserfields": { + "exclouded": ["id"], + "cwidth": {} + }, + "editexclouded": ["id"], + "subtables": [ + { + "field": "accountid", + "title": "账户余额", + "subtable": "acc_balance" + }, + { + "field": "accountid", + "title": "账户明细", + "subtable": "acc_detail" + }, + { + "field": "accountid", + "title": "账户日志", + "subtable": "accounting_log" + } + ] + } +} +``` + +### 3. Frontend Integration +- **UI Files**: `.ui` files in wwwroot/ use Jinja2 templating +- **Script Files**: `.dspy` files provide server-side logic +- **Assets**: Static resources in wwwroot/imgs/ + +### 4. Business Logic Organization +Core functionality is organized into logical modules: +- `accounting_config.py`: Configuration management +- `bill.py`: Billing operations +- `openaccount.py`: Account creation workflows +- `getaccount.py`: Account querying +- `recharge.py`: Recharge processing +- `consume.py`: Consumption processing +- `ledger.py`: Ledger operations + +## Compliance Verification + +### ✅ Module Development Specification Compliance +- [x] Proper directory structure with accounting/, wwwroot/, json/, models/ +- [x] Correct init.py with load_accounting() function +- [x] ServerEnv exposure of all required functions +- [x] CRUD definitions in json/ directory +- [x] Frontend resources in wwwroot/ directory +- [x] Database table definitions in models/ directory + +### ⚠️ Minor Deviations +- Uses `.xlsx` format for table definitions instead of `.json` (still valid internal format) +- Uses `setup.py` instead of `pyproject.toml` (legacy but functional) +- Missing `init/data.json` (optional if no initialization data needed) + +## Usage as Reference Implementation + +This accounting module serves as an excellent reference for: +1. **Module Structure**: How to organize a complex business module +2. **Function Exposure**: Proper ServerEnv usage patterns +3. **CRUD Configuration**: Real-world CRUD definition examples +4. **Business Logic**: Separation of concerns in accounting operations +5. **Frontend Integration**: UI/script resource organization + +## Deep-Dive References + +- **`references/accounting-internals.md`** — PFBiz/Accounting class hierarchy, leg_accounting hot path, overdraft check pattern, credit limit extension, subject/account relationships, accounting_config table driving journal entries +- **`references/coupon-system.md`** — platformbiz coupon/coupontype/coupon_log table structure, mintransamt (满减门槛) support, gap analysis for tiered discounts, integration plan with accounting module +- **`references/credit-limit-multi-tenant.md`** — multi-tenant credit limit design: grant_orgid field, admin vs customer read views, migration SQL + +## Pitfalls + +### Balance update is NOT optional +Accounting = 写分录明细 + 写日志 + 修改账户余额. These three steps are the DEFINITION of accounting (记账), not optional add-ons. If an implementation only writes the detail record without updating the balance, it is incomplete by definition — do not characterize balance update as a "missing feature" or "nice to have". It IS the accounting. + +### Overdraft check belongs in the same transaction +The balance update and overdraft/credit-limit check must happen in the same DB context as the detail insert. Reading balance, checking credit limit, and writing the new balance must be atomic with the accounting record. + +## Integration Notes +- Integrates with sqlor-database-module for database operations +- Uses bricks-framework compatible UI templates +- Follows security patterns from user/org context handling +- Implements comprehensive accounting workflows (recharge, consume, billing, balance queries) + +## Integration Checklist for New Features + +When adding any new entity/feature to the accounting module (or any Sage module), you MUST verify all four integration points: + +1. **init.py** — New functions/classes must be imported and exposed via `ServerEnv` in `load_()` +2. **scripts/load_path.py** — All new `.ui` and `.dspy` pages must have RBAC paths registered (in module's own `scripts/` directory, not sage main repo) +3. **wwwroot/global_menu.ui** (sage main repo) — Menu entry for the new page +4. **json/.json** — CRUD definition file (must conform to crud-definition-spec: root keys = tblname + params) + +Missing any of these means the feature is invisible/unusable even if the code is correct. Always audit all four before declaring done. + +## Pitfalls + +- **Don't revert approved changes when context expands.** If the user approves change A, and later says "also do B", that doesn't mean A was wrong. Build on A, don't undo it. User frustration: "为什么实际做却不按确认的做呢" — reverting approved work without asking. +- **Read the full existing module before modifying.** The accounting module has a complete system (PFBiz → Accounting → leg_accounting). Before adding features, read `accounting_config.py`, `creditlimit.py`, `consume.py`, etc. to understand how they work together. Don't invent parallel implementations. +- **sageapi vs sage accounting are two layers.** sage/pkgs/accounting/ is the core accounting engine (double-entry, legs, subjects). sageapi is a lightweight API gateway. Both may need credit_limit logic but in different ways — don't confuse them. + +## Learning Points +- How to expose both classes and functions through ServerEnv +- Pattern for async database query functions with proper context management +- Subtable relationships in CRUD definitions for master-detail scenarios +- Organization of complex business logic across multiple Python modules \ No newline at end of file diff --git a/skills_library/all/agentic-report-generation/SKILL.md b/skills_library/all/agentic-report-generation/SKILL.md new file mode 100644 index 0000000..1169715 --- /dev/null +++ b/skills_library/all/agentic-report-generation/SKILL.md @@ -0,0 +1,133 @@ +--- +name: agentic-report-generation +description: "Agentic report generation with per-step quality gates." +tags: [agentic, report-generation, llm, evaluation, hitl, quality-gates, sage] +triggers: + - "agentic report generation" + - "report generation pipeline" + - "quality-gated generation" + - "分步生成" + - "报告生成" + - "auto-evaluate report" + - "尽调报告" +--- + +# Agentic Report Generation (Quality-Gated Pipeline) + +How to design an agentic multi-section report/document generation system where each section is independently generated, auto-evaluated against a quality threshold, retried with feedback, escalated to a human only when retries are exhausted, then merged into a deliverable. + +## When to Use + +- User asks to design/build a system that generates a multi-section report (尽调报告, financial analysis, audit report, assessment, etc.) from source materials. +- Need to guarantee report quality is *predictable*, not luck-of-the-draw. +- Need customer-customizable output format (their own template) and metric definitions (their own indicator set). + +## Core Pipeline Pattern + +Do NOT generate the whole report in one LLM call. Split into independent section-steps, each with its own quality gate: + +``` +[生成] 分步骤生成各章节(各自独立) + ↓ +[评估] 自动三维打分(完整性 / 准确性 / 合规性) + ↓ + ├─ 评分 ≥ 阈值(默认80分)──→ 步骤达标,进入下一步 + ├─ 评分 < 阈值,重试 < N次(默认3次)──→ 参照评估建议自动重生成 + └─ 重试 ≥ N次 仍不达标 ──→ 转人工干预 + ↓ +[合并] 各步骤全部达标 → 自动合并成交付文档(套用客户模板) +``` + +## Auto-Evaluation Dimensions + +Score each step on three axes (produce a total + per-axis breakdown): + +| 维度 | 评估内容 | +|------|---------| +| 完整性 completeness | 该章节必填字段/要素是否齐全(对照模板章节定义) | +| 准确性 accuracy | 数据是否与源材料一致(溯源校验),有无杜撰、遗漏 | +| 合规性 compliance | 格式是否符合模板、结论是否有依据、是否覆盖框架要点 | + +The evaluator must output **specific 评估建议 (advice)** — a concrete problem list (e.g. "缺少抵押物查封顺位信息", "估值折扣率未说明依据") — not just a score. The generator consumes this advice on retry. + +## Threshold + Retry + HITL + +- **Threshold**: per-step pass line, default 80, configurable. +- **Retry**: generator re-runs with the advice appended, up to N times (default 3). +- **Human-in-the-loop (HITL)**: after N failed retries, escalate to the user. The user views the problem list + provides a **natural-language instruction** ("补充查封顺位为第一顺位", "将折扣率调整为65%") → regenerate that step with the instruction. The user may also hand-edit the section directly. + +## Customization-as-Config (客制化) Principle + +For customer-facing systems, do NOT hardcode the two things customers always want to own: + +1. **Metric/indicator definitions** → an editable **indicator tree** (可视化树形维护): each node = {name, definition, formula, data_source, threshold}. Customer adds/edits/deletes nodes; report generation reads the tree to compute the indicator summary. Default tree preset, customer customizes on top. + + **Three-level hierarchy (三级客制化)**: customers usually want per-organization AND per-project customization. Model it as three scopes that inherit downward, each gated by approval: + - `system` 系统缺省 — built-in default, read-only. + - `company` 公司通用 — customer customizes the default; on approval becomes the company-wide standard. + - `project` 项目专用 — a single project further customizes company scope; on approval applies to that project only. + Resolution order at report time: **project > company > system**. Add `scope` + `asset_id` + `approval_status` to the node table. +2. **Output format** → a **customer-provided template** (upload Word/Excel, system parses placeholders + section structure, fills on merge). Multiple templates coexist (report / finance / valuation), selected per task. + + **Dual-version output**: client-facing reports often need BOTH a Word version (detailed — internal review / archiving / signing) and a PPT version (presentation — management decision / external communication). Generate content ONCE, then render through two independent templates. Add a `format` field (word / ppt) to the template so both versions coexist per report type. + +## Data Model + +```sql +-- report step (one per section) +CREATE TABLE report_step ( + id VARCHAR(32) PRIMARY KEY, + report_id VARCHAR(32), + step_no INT, -- S1..Sn + section_name VARCHAR(200), + content TEXT, -- step output + score FLOAT, -- latest quality score + status VARCHAR(20), -- pending/generating/passed/retrying/manual + retry_count INT DEFAULT 0, + max_retry INT DEFAULT 3, + pass_threshold FLOAT DEFAULT 80 +); + +-- per-step evaluation record (auto + manual) +CREATE TABLE report_evaluation ( + id VARCHAR(32) PRIMARY KEY, + step_id VARCHAR(32), + completeness_score FLOAT, + accuracy_score FLOAT, + compliance_score FLOAT, + total_score FLOAT, + advice TEXT, -- 评估建议(问题清单) + passed TINYINT, + eval_type VARCHAR(20), -- auto / manual + user_instruction TEXT, -- 人工干预指令(manual 时) + created_at DATETIME +); + +-- indicator tree (customer-editable, three-level scope) +CREATE TABLE indicator_node ( + id VARCHAR(32) PRIMARY KEY, + scope VARCHAR(20), -- system / company / project + asset_id VARCHAR(32), -- set when scope=project + parent_id VARCHAR(32), + name VARCHAR(200), + node_type VARCHAR(20), -- category / indicator + definition TEXT, formula TEXT, data_source VARCHAR(64), threshold TEXT, + approval_status VARCHAR(20), -- pending / approved (company+project need approval) + sort_order INT, is_active TINYINT +); + +-- report template (customer-provided, per output format) +CREATE TABLE report_template ( + id VARCHAR(32) PRIMARY KEY, + name VARCHAR(200), type VARCHAR(50), + format VARCHAR(10), -- word / ppt + file_path VARCHAR(500), structure TEXT, is_default TINYINT +); +``` + +## Key Pitfalls + +- **Never generate the whole report in one call** — a single bad section forces regenerating everything, and you can't measure per-section quality. +- **Evaluation must emit advice, not just a score** — a bare "72分" gives the generator nothing to fix on retry. +- **Traceability**: every report conclusion should carry a source-material reference, so "准确性" evaluation and downstream audit can verify it. +- **Customization is a selling point, not an afterthought** — build the indicator tree and template upload as first-class features from day one (customers reject hardcoded metric/format systems). diff --git a/skills_library/all/ahserver-hot-reload/SKILL.md b/skills_library/all/ahserver-hot-reload/SKILL.md new file mode 100644 index 0000000..d013aff --- /dev/null +++ b/skills_library/all/ahserver-hot-reload/SKILL.md @@ -0,0 +1,449 @@ +--- +name: ahserver-hot-reload +version: 1.0.0 +description: ahserver file-based hot-reload system for config, i18n, and module caches (multi-process safe) +trigger_conditions: + - User asks how to enable/configure ahserver hot-reload + - User needs to clear module caches without restarting + - User is debugging stale config/i18n/cache in multi-worker deployment + - User asks about /__hot_reload__ endpoint + - User needs to distribute cache invalidation across multiple workers +--- + +# ahserver Hot-Reload System + +File-based hot-reload for ahserver, watching config.json and i18n files. Multi-process safe (each worker independently checks file mtimes). + +Code: `ahserver/ahserver/hotreload.py`, `ahserver/ahserver/webapp.py` + +## Enable Hot-Reload + +Add to `conf/config.json`: + +```json +{ + "hot_reload": true +} +``` + +Or with custom interval: + +```json +{ + "hot_reload": { + "enabled": true, + "interval": 2 + } +} +``` + +Restart ahserver after enabling. + +## What Gets Auto-Reloaded + +| Trigger | Config Singleton | Module Caches (hot_reload event) | +|---------|-----------------|----------------------------------| +| `conf/config.json` mtime change | ✓ Cleared (next getConfig reloads) | ✗ NOT dispatched | +| `i18n/*/msg.txt` mtime change | ✗ | ✓ Dispatched | +| Signal file mtime change (cross-worker) | ✗ | ✓ Dispatched | +| `GET /__hot_reload__` endpoint | ✗ | ✓ Dispatched (also writes signal file) | + +**Key design**: config.json changes only refresh the JsonConfig singleton. Module caches are NOT cleared because config changes rarely affect cached module data. Only i18n changes, signal file updates, or explicit HTTP calls trigger cache clearing. + +## Manual Cache Invalidation + +HTTP endpoint to clear all module caches without file changes: + +```bash +curl http://localhost:PORT/__hot_reload__ +``` + +Returns: +```json +{ + "status": "ok", + "message": "Signal sent to all workers, current worker dispatched hot_reload", + "timestamp": 1738416000.0 +} +``` + +## Module Caches Cleared (via EventDispatcher) + +Each module implements `on_hot_reload(data=None)` on its cache-holding class/instance, bound in `load_XXX()`: + +| Module | Cache | Clear Method | Bound On | +|--------|-------|--------------|----------| +| rbac | User permissions (LRUCache), role-permissions (dict→None) | `UserPermissions.on_hot_reload()` | Instance method (stored on ServerEnv) | +| pricing | Pricing data per org (class-level dict) | `PricingProgram.on_hot_reload()` | @staticmethod (class-level) | +| uapi | API users, API definitions, API keys (3 dicts) | `UAPIData.on_hot_reload()` | Instance method (stored on ServerEnv) | +| llmage | LLM API/uapiio cache (module-level dicts) | `_on_hot_reload()` wrapper | Module-level function (module keeps it alive) | + +### ⚠️ CRITICAL: rbac is NOT a singleton + +`UserPermissions` does NOT use `@SingletonDecorator`. The actual instance is created once in `rbac/load_rbac()` and stored on `ServerEnv().userpermissions`. Creating `UserPermissions()` anywhere else gives you a **new empty instance** with empty caches — clearing it does nothing. + +**Wrong**: +```python +from rbac.userperm import UserPermissions +up = UserPermissions() # ← NEW empty instance, not the real one +up.ur_caches.clear() # ← clears nothing useful +``` + +**Correct**: +```python +from ahserver.serverenv import ServerEnv +g = ServerEnv() +up = g.userpermissions # ← the actual instance with real caches +up.ur_caches.clear() +up.invalidate_rp_cache() +``` + +## Multi-Worker Deployment + +ahserver runs multiple workers (via `reuse_port=True`). Each worker has independent Python memory space. + +### Problem + +`GET /__hot_reload__` only clears cache in the worker that receives the request. Other workers still have stale cache. + +### Solutions (choose based on deployment) + +#### Solution 1: Shell Loop (simplest, reuse_port multi-port) + +Each worker listens on different port: + +```bash +#!/bin/bash +# hot_reload_all.sh +PORTS=(8000 8001 8002 8003) +for port in "${PORTS[@]}"; do + curl -s "http://127.0.0.1:$port/__hot_reload__" & +done +wait +echo "Done" +``` + +**When to use**: reuse_port mode with explicit port assignment. + +Ready-to-use script: `scripts/hot_reload_all.sh` — pass ports as args or defaults to 8000-8003. + +#### Solution 2: nginx mirror directive (automatic replication) + +```nginx +upstream worker_0 { server 127.0.0.1:8000; } +upstream worker_1 { server 127.0.0.1:8001; } +upstream worker_2 { server 127.0.0.1:8002; } +upstream worker_3 { server 127.0.0.1:8003; } + +server { + listen 80; + + location /__hot_reload__ { + mirror /__hot_reload_mirror_1__; + mirror /__hot_reload_mirror_2__; + mirror /__hot_reload_mirror_3__; + + proxy_pass http://worker_0; + } + + location = /__hot_reload_mirror_1__ { + internal; + proxy_pass http://worker_1/__hot_reload__; + } + location = /__hot_reload_mirror_2__ { + internal; + proxy_pass http://worker_2/__hot_reload__; + } + location = /__hot_reload_mirror_3__ { + internal; + proxy_pass http://worker_3/__hot_reload__; + } +} +``` + +**When to use**: nginx as load balancer, want single curl to hit all workers. + +**Pitfall**: nginx mirror is fire-and-forget — client doesn't see mirror responses. If a worker fails, you won't know from the main response. + +#### Solution 3: File Signal (IMPLEMENTED — production default) + +Already implemented in `hotreload.py` (commit 42eff6c). No code changes needed. + +**How it works**: +1. `GET /__hot_reload__` hits any worker via nginx +2. That worker writes timestamp to `/tmp/.sage_cache_invalidate` and dispatches `hot_reload` immediately +3. All other workers' `HotReloader._check_signal_file()` detects mtime change within `interval` seconds (default 2s) +4. All workers dispatch `hot_reload` event → each module's bound handler clears its own cache +5. Each worker clears its own caches independently + +**Single curl is sufficient** — no shell loop or nginx config needed: +```bash +curl http://localhost:PORT/__hot_reload__ +``` + +Response only shows the worker that received the request, but ALL workers will clear caches within ~2s. + +## Logs + +**INFO level** (default): +``` +[hot_reload] started, interval=2s +[hot_reload] reloaded: ['config', 'i18n'] +[hot_reload] reloaded: ['signal'] +[hot_reload] stopped +``` + +**DEBUG level** (set `logger.levelname: "debug"` in config.json): +``` +[hot_reload] config_path=/path/to/conf/config.json +[hot_reload] watching 2 i18n paths +[hot_reload] initial mtime for /path/to/file: 1717257600.0 +[hot_reload] changed: /path/to/file (mtime 1717257600.0 -> 1717257700.0) +[hot_reload] signal file mtime: 1717257700.0, last: 0 +[hot_reload] signal file changed, triggering reload +[hot_reload] config changed: ['/path/to/conf/config.json'] +[hot_reload] clearing JsonConfig singleton +[hot_reload] clearing MiniI18N singleton +[hot_reload] cleared ServerEnv.myi18n +[hot_reload] config-only change, skipping cache clear dispatch +[hot_reload] dispatching hot_reload event (non-config changes detected) +[hot_reload] HTTP endpoint triggered, writing signal to /tmp/.sage_cache_invalidate +[hot_reload] HTTP endpoint: dispatching hot_reload event +``` + +**Module handler logs** (DEBUG level): +``` +[uapi] on_hot_reload called, clearing caches (data={...}) +[rbac] on_hot_reload called, clearing caches (data={...}) +[pricing] on_hot_reload called, clearing pricing_data (data={...}) +[llmage] on_hot_reload called, invalidating uapi cache (data={...}) +``` + +**Troubleshooting**: If hot_reload isn't triggering cache clears, enable DEBUG logging and check: +1. File mtime changes are detected (look for `changed:` log) +2. Whether it's config-only (look for `config-only change, skipping` vs `dispatching`) +3. Whether module handlers are called (look for `on_hot_reload called` logs) +4. If handler logs missing, check the module's `load_XXX()` bind call — WeakCallback may have lost the reference + +## Limitations + +1. **No Python code hot-reload** — Only config/i18n/cache. Code changes require restart. +2. **File mtime resolution** — On some filesystems (NFS, Docker volumes), mtime may not update immediately. +3. **Signal file latency** — Multi-worker cache clear has ~2s delay (configurable via `interval`). Not instant like Redis Pub/Sub would be. + +## Comparison with Redis Pub/Sub cache_sync + +| Feature | hot-reload (this) | cache_sync (Redis) | +|---------|-------------------|-------------------| +| Trigger | File change / HTTP | Database event | +| Infrastructure | None | Redis | +| Latency | 2s (polling) | Instant | +| Status | **Production ready** | Reverted (session loss bug) | +| Use case | Dev/testing, manual invalidation | Production auto-sync | + +See `sage-cache-sync` skill for Redis Pub/Sub approach (currently reverted). + +## EventDispatcher Architecture (Implemented) + +Uses `appPublic.event_dispatcher.EventDispatcher` (NOT `eventpy`) — implements WeakCallback with weakref for automatic cleanup. See `references/event-dispatcher-api.md` for full API reference. + +### Key API + +```python +class EventDispatcher: + def bind(self, event_name: str, func: Callable) # register handler (WeakCallback) + def unbind(self, event_name: str, func: Callable) # unregister + async def dispatch(self, event_name: str, data=None) # fire event, await all handlers +``` + +Handlers receive `data` argument (the reloaded dict or custom payload). Both sync and async handlers are supported. + +### Lifecycle + +``` +webserver() in webapp.py: + 1. se.event_dispatcher = EventDispatcher() ← BEFORE init_func() + 2. init_func() → load_rbac/pricing/uapi/llmage → each binds 'hot_reload' + 3. ConfiguredServer → server.run() + +Runtime triggers → dispatch('hot_reload'): + - GET /__hot_reload__ → writes signal file + immediate dispatch + - signal file mtime change (other workers) → dispatch + - i18n file mtime change → dispatch + +Config.json mtime change → reloads JsonConfig singleton ONLY, does NOT dispatch hot_reload. +``` + +### Adding a New Module's Cache Clear + +In your module's class, add `on_hot_reload`: + +```python +class MyModule: + def __init__(self): + self.cache = {} + + def on_hot_reload(self, data=None): + self.cache.clear() +``` + +In `load_mymodule()`: + +```python +def load_mymodule(): + env = ServerEnv() + env.mymodule = MyModule() + # Guard for non-web contexts (scripts, tests) + # CRITICAL: use getattr + None check, NOT hasattr + # hasattr only checks attribute existence, but event_dispatcher + # can exist as None when running standalone (e.g. backend_accounting.py) + if getattr(env, 'event_dispatcher', None) is not None: + env.event_dispatcher.bind('hot_reload', env.mymodule.on_hot_reload) +``` + +### ⚠️ CRITICAL: WeakCallback Pitfalls + +EventDispatcher uses `weakref.ref` for functions and `weakref.WeakMethod` for instance methods. If the handler's target gets garbage-collected, the binding silently disappears. + +**Wrong — lambda gets GC'd immediately:** +```python +env.event_dispatcher.bind('hot_reload', lambda data: cache.clear()) +# lambda has no strong reference → GC'd → binding lost +``` + +**Wrong — local function gets GC'd:** +```python +def load_mymodule(): + async def clear(data): # local function + cache.clear() + env.event_dispatcher.bind('hot_reload', clear) + # clear() is local → GC'd after load_mymodule() returns → binding lost +``` + +**Correct patterns:** + +| Pattern | Why it works | +|---------|-------------| +| Instance method on object stored on `ServerEnv` | `ServerEnv` holds strong ref to instance → WeakMethod stays valid | +| `@staticmethod` on a class | Class is never GC'd → ref stays valid | +| Module-level function | Module stays loaded → ref stays valid | + +### Signature Requirement + +All handlers receive `data` as argument. If wrapping an existing function that doesn't accept args: + +```python +# llmage's invalidate_uapi_cache() takes optional upappid/apiname +# dispatcher calls with data=dict → need wrapper +def _on_hot_reload(data=None): + invalidate_uapi_cache() + +env.event_dispatcher.bind('hot_reload', _on_hot_reload) +``` + +## Pitfalls + +### rbac UserPermissions is not a singleton + +`UserPermissions()` creates a new empty instance. Always use `ServerEnv().userpermissions` to get the actual instance with real caches. See: +- `references/rbac-non-singleton-pitfall.md` — why this happens and how to avoid it +- `references/rbac-event-handler-bug.md` — unfixed bug in rbac/init.py event handlers (same root cause) + +### hasattr vs getattr for event_dispatcher — use getattr with None check + +`hasattr(env, 'event_dispatcher')` only checks attribute existence. In standalone scripts (e.g., `backend_accounting.py`), `event_dispatcher` exists on `ServerEnv` but its value is `None`. This causes `AttributeError: 'NoneType' object has no attribute 'bind'`. + +**Wrong:** +```python +if hasattr(env, 'event_dispatcher'): + env.event_dispatcher.bind('hot_reload', handler) # ← crashes if event_dispatcher is None +``` + +**Correct:** +```python +if getattr(env, 'event_dispatcher', None) is not None: + env.event_dispatcher.bind('hot_reload', handler) +``` + +### Debug log noise in periodic tasks + +The hot_reload task runs every N seconds and checks multiple file mtimes. Debug logs that fire unconditionally on every check cycle flood the log file and obscure real events. + +**Wrong — logs every 2s even when nothing changes:** +```python +def _check_signal_file(self): + mtime = os.path.getmtime(SIGNAL_FILE) + debug(f'[hot_reload] signal file mtime: {mtime}, last: {self._last_signal_mtime}') # ← noise + if mtime > self._last_signal_mtime: + ... +``` + +**Correct — only log when state actually changes:** +```python +def _check_signal_file(self): + mtime = os.path.getmtime(SIGNAL_FILE) + if mtime > self._last_signal_mtime: + self._last_signal_mtime = mtime + debug(f'[hot_reload] signal file changed, mtime: {mtime}') # ← only on change + return True +``` + +**Same applies to OSError on missing files** — the signal file may not exist for hours. Don't log "not found" on every check; silently pass. + +**General rule for periodic task debug logging**: Gate log statements behind the condition that makes them interesting. "Checked X" is noise; "X changed from A to B" is signal. + +### Config.json must be valid JSON + +Hot-reload clears JsonConfig singleton, next `getConfig()` reloads from disk. If config.json has syntax error, server will crash on next config access. + +**Fix**: Validate config.json before saving. + +### aiohttp cleanup_ctx vs on_cleanup + +`app.cleanup_ctx.append()` requires an **async context manager** (must `yield`). Plain `async def` functions that don't yield cause `AttributeError: 'coroutine' object has no attribute '__aiter__'`. + +| API | Accepts | Use for | +|-----|---------|---------| +| `app.cleanup_ctx.append()` | `async def f(app): ... yield ...` (async context manager) | Need setup + teardown in one function | +| `app.on_cleanup.append()` | `async def f(app): ...` (plain coroutine) | Teardown-only cleanup (e.g. cancel task) | + +**Bug in hot_reload**: `_hot_reload_cleanup` was a plain async def added to `cleanup_ctx`. Fixed by switching to `on_cleanup.append()`. + +**Symptom**: +``` +AttributeError: 'coroutine' object has no attribute '__aiter__'. Did you mean: '__dir__'? +sys:1: RuntimeWarning: coroutine '_hot_reload_cleanup' was never awaited +``` + +### i18n file path detection + +`get_i18n_paths()` scans `i18n/*/msg.txt`. If you add a new language directory after hot-reload starts, it won't be watched until restart. + +**Fix**: Restart after adding new language. + +### Signal file detection + +Signal file at `/tmp/.sage_cache_invalidate` — all workers detect mtime change and dispatch `hot_reload` event. + +### Module import errors (no longer applies) + +Previously `invalidate_all_caches()` imported modules directly. Now uses EventDispatcher — modules self-register. If a module doesn't bind, its cache won't be cleared (check its `load_XXX()` for the bind call). + +### Git force-commit needed to resync truncated files + +When a file is truncated in the server's working copy but the local repo already has the correct version at HEAD, `git checkout HEAD` reports no change and `git pull` says "up to date". The server never gets the fix. + +**Symptom**: Server returns 500 because a function is missing from a truncated file, but `git pull` on server shows nothing to update. + +**Root cause**: The file was modified locally (truncated), committed, then restored via `git checkout HEAD`. Now local and remote HEAD are identical — the correct file is in git history. But the server's working copy still has the old truncated version. + +**Fix**: Force a commit that changes the file, even trivially: +```bash +# Add a comment or whitespace to create a diff +echo "# Force re-sync" >> path/to/file.py +git add path/to/file.py +git commit -m "force: re-sync (ensure full version)" +git push +``` + +Then server `git pull` will pull the new commit and overwrite the truncated file. diff --git a/skills_library/all/ahserver-pitfalls/SKILL.md b/skills_library/all/ahserver-pitfalls/SKILL.md new file mode 100644 index 0000000..d65f16f --- /dev/null +++ b/skills_library/all/ahserver-pitfalls/SKILL.md @@ -0,0 +1,318 @@ +--- +name: ahserver-pitfalls +description: "POST 405 fix, auth crash, multipart hang. Voiceprint howto." +version: "1.0.0" +--- +# ahserver Pitfalls & Voiceprint Integration + +## POST 405 for startswiths Routes +aiohttp StaticResource tracks allowed methods in `_allowed_methods` set. `ProcessorResource.__init__` adds POST to `_routes` but not `_allowed_methods`, so POST returns 405 with `Allow: GET,HEAD`. + +**Fix:** In `processorResource.py` `__init__`, after `_routes.update` lines: +```python +self._allowed_methods = set(self._routes.keys()) +``` + +## Auth Middleware Crash on Multipart Uploads +`get_session_userinfo` calls `auth.get_auth(request)` which raises `RuntimeError('auth_middleware not installed')` when auth is disabled via `self.user = None`. This crashes `getPostData` during multipart processing. + +**Fix in `auth_api.py`:** +```python +async def get_session_userinfo(request): + try: + d = await auth.get_auth(request) + except: + d = None + if d is None: + return DictObject() +``` + +## client_max_size Too Small → Silent Hang +`conf/config.json` `client_max_size: 10000` (10KB) causes large multipart uploads to hang. Small files work, large files (>client_max_size) never return a response. + +**Fix:** Set to >= expected max file size. For audio/video: `104857600` (100MB). + +## Voiceprint Service (media.opencomputing.net:10443) +- Location: `ymq@opencomputing.net:/share/ymq/run/voiceprint` +- Start: `PYTHONPATH='.:sqlor:ahserver:appPublic:longtasks' python3 ah.py -p 9087` +- GPU: cuda:1, ECAPA-TDNN model via speechbrain +- Endpoint: `POST /extract/submit` with multipart `file` field +- Response: `{"status":"SUCCEEDED","embedding":[...],"embedding_dim":192}` — no `speakers` field + +## speechbrain load_audio Signature +`SpeakerRecognition.load_audio(self, path, savedir=None)` — 2nd arg is `savedir`, NOT sample rate. +- ❌ `load_audio(path, 16000)` → TypeError (int as Path) +- ✅ `load_audio(path)` + +## sqlor `IN (${ids}$)` List Expansion Failure +On some sqlor versions, passing a Python list to `${ids}$` for `IN` clauses raises: +``` +Illegal parameter data types varchar and row for operation '=' +``` +**Workaround** — build quoted comma-separated string manually: +```python +id_list = ','.join(["'" + str(x) + "'" for x in doc_ids]) +# Then use in raw SQL concatenation: +"... WHERE id IN (" + id_list + ")" +``` +Always wrap in `try/except` as fallback. + +## DSPY `%%` LIKE Patterns — Avoid Escaped Quotes +Using `\"` inside a `%%...%%` LIKE pattern in a double-quoted Python string causes SyntaxError: +```python +# ❌ BROKEN — \" closes the Python string +"... AND metadata LIKE '%%voiceprint_status%%\"done\"%%' ..." + +# ✅ Use simple patterns without quoted substrings: +"... AND metadata LIKE '%%voiceprint_status%%done%%' ..." +``` + +## SSH Background Process on This Server +`nohup ... &` hangs SSH. Preferred order: +1. `ssh -f user@host "cmd"` — forks background, returns immediately +2. Use `terminal(background=true)` — Hermes's own background mode + +## Multipart Handler Pattern +ahserver auto-handles multipart: file saved to FileStorage, params_kw has web_path: +```python +web_path = params_kw.get('file') +fs = FileStorage() +abs_path = fs.realPath(web_path) +``` + +### Frontend: bricks 上传文件必须用 FormData,JSON.stringify 会静默丢 File 对象 +bricks `UiFile` 只把浏览器 `File` 对象存进 `this.value`(内存),**不会自动上传**。`AgentIO`/`TextFiles` 经 `HttpResponseStream.post → bricks_fetch` 发送,当 params 不是 FormData 时走 `JSON.stringify(data)` —— **File 对象被序列化成 `{}`,文件内容静默丢失**(只有 `f.name` 字符串能传出去)。这就是"用户上传了文件但后端 agent 收不到内容"的根因。 + +修复(`bricks/agent.js` 的 `user_inputed`):有 `add_files` 时构造 FormData 上传文件二进制: +```javascript +var files = params.add_files || []; +var send_params = params; +if (files.length > 0) { + send_params = new FormData(); + Object.keys(params).forEach(function(k){ + if (k !== 'add_files' && k !== 'file_names') send_params.append(k, params[k]); + }); + files.forEach(function(f){ send_params.append('file', f); }); +} +var resp = await hr.post(this.opts.url, {params:send_params}); +``` +`bricks_fetch` 已处理 `data instanceof FormData`(自动 append session、body=FormData)。改完需重新 build `dist/bricks.js`:`bash build.sh`(把 `bricks/*.js` 按 SOURCES 列表 cat 合并到 dist,前端加载的是 dist 打包版,不是源码)。 + +dspy 后端取文件:`params_kw.get('file')`(单个 web_path 或 list),`FileStorage().realPath()` 拿绝对路径。docx 文本提取:zipfile 读 `word/document.xml` + `re.findall(r']*>(.*?)', xml)`(`cat` 读 docx 是乱码,必须解 zip 提取 ``)。 + +## DSPY Silent Error Swallowing +DSPY files often wrap logic in `except Exception: return "加载失败"`. This silently hides the real error. When debugging, always replace with: +```python +except Exception as _e: + import traceback + return {"widgettype":"Text","options":{"text":str(_e)+"\n"+traceback.format_exc()[-200:]}}} +``` +Common hidden errors: sqlor placeholder mismatch, Python SyntaxError in string concatenation, missing imports. + +## Tag Storage: Dual Sources (media_tags table + metadata.tags) +Tags can live in TWO places: +1. `media_tags` + `tags` tables — from face processing, tag assignment +2. `documents.metadata.tags` JSON array — from `add_tag.dspy` UI + +When displaying tags on cards, read from BOTH sources: +```python +# Source 1: media_tags table +doc_tags = {} +try: + id_list = ','.join(["'" + str(x) + "'" for x in doc_ids]) + mt_recs = await sor.sqlExe( + "SELECT mt.media_id, t.name, t.color FROM media_tags mt " + + "JOIN tags t ON mt.tag_id=t.id " + + "WHERE mt.media_type='document' AND mt.media_id IN (" + id_list + ")", + ns={}) + for mt in mt_recs: + doc_tags.setdefault(mt.media_id, []).append({"name": mt.name, "color": mt.color}) +except: pass +# Source 2: metadata.tags from add_tag.dspy +for r in rows: + try: + meta = json.loads(r.get("metadata", "{}")) + for t in meta.get("tags", []): + # deduplicate + existing = doc_tags.get(r["id"], []) + if not any(e.get("name")==t for e in existing): + existing.append({"name": t, "color": "#3b82f6"}) + doc_tags[r["id"]] = existing + except: pass +``` + +## Media URLs in DSPY: Use entire_url() +Widgets like VideoPlayer/Image/Audio need absolute URLs. `safe_url()` only prepends `/idfile`: +```python +# ❌ Relative path — breaks in some contexts +media_url = safe_url(h.get("file_path", "")) +# ✅ Absolute URL +media_url = entire_url(safe_url(h.get("file_path", ""))) +# Result: https://rag.opencomputing.cn/idfile/117/169/.../file.mp4 +``` + +## Media Cards: Voice Query Including Videos +Voice cards should show BOTH audio files AND videos with extracted voiceprints: +```python +# Query includes videos that have voiceprint_status=done +"WHERE kb_id=${kb_id}$ AND (metadata LIKE '%%voiceprint_status%%done%%' " + +"OR LOWER(file_name) LIKE '%%.mp3' OR LOWER(file_name) LIKE '%%.wav' ...)" +``` + +## Video Playback: Native `
.json vs hand-written wwwroot/
_list/ + +When `json/
.json` (CRUD DataViewer config) exists, the Sage framework auto-generates CRUD endpoints for `/
_list/`. Creating hand-written `wwwroot/
_list/` with custom DSPY files (data.dspy, add.dspy, etc.) conflicts with the auto-generated routes and causes **403 Forbidden**. + +**Fix**: If `json/
.json` exists, delete the hand-written `wwwroot/
_list/` directory. Let the framework auto-CRUD handle the list page. The `json/
.json` editfields/browserfields control the Tabular widget. + +### 🔴 upload_file.dspy deployed WITHOUT PDF/DOCX/PPTX/XLSX extraction → silent zero-chunk + +The deployed `upload_file.dspy` can be MISSING the `elif` branches for PDF/DOCX/PPTX/XLSX text extraction — only has the plain-text `if ext_l in text_exts` branch. When non-text files (PDF, DOCX, PPTX, MP4, etc.) are uploaded, `text` stays `''`, the `if text and len(text.strip()) > 10:` guard is never true, and the entire RAG ingest block (chunking → embedding → VDB) is skipped. Documents get inserted with `status='done'` and `chunk_count=0` — **looks successful, nothing ingested**. + +The correct full pattern (with `elif ext_l == '.pdf': ... elif ext_l == '.docx': ...` etc.) IS in `references/rag-ingest-pipeline.md`. Before deploying, diff against that reference. + +**Symptom**: `document_chunks` is empty for all non-plain-text files. + +**Verification query** (run after any upload deploy): +```bash +mysql -h db -u test -ptest123 rag -e " +SELECT d.file_name, d.status, COUNT(c.id) AS chunks +FROM documents d LEFT JOIN document_chunks c ON c.doc_id = d.id +WHERE d.kb_id = '' +GROUP BY d.id, d.file_name, d.status +HAVING chunks = 0" +``` + +Any row in the output = a file that uploaded but was never ingested. + +### 🔴 base64 NOT in DSPY pre-loaded globals — crashes image/video upload + +`base64` module is **NOT** among ahserver's DSPY pre-loaded globals (only `hex2base64` function is). Using `base64.b64encode()` without `import base64` raises `NameError: name 'base64' is not defined`, crashing the entire DSPY before the DB insert. + +Confirmed in server logs: +``` +NameError: name 'base64' is not defined + at upload_file.dspy line 41: img_b64 = base64.b64encode(file_data).decode() +``` + +**Fix**: Add `import base64` at the top of any DSPY file that uses base64 (e.g., for face detection on image/video frames to build data URIs for GPU API calls). + +### 🔴 DSPY HTTP calls: use StreamHttpClient, NOT aiohttp + +`StreamHttpClient` IS a pre-loaded DSPY global (from `globalEnv.py`: `g.StreamHttpClient = StreamHttpClient`). `aiohttp` is NOT available in DSPY context. Using `aiohttp.ClientSession(...)` without import raises `NameError`. The critical danger: if wrapped in bare `except: pass`, the NameError is silently swallowed — the file upload succeeds (DB insert at end runs), but EVERY processing step fails silently: no embedding, no face detection, no voiceprint, no VDB upsert. Result: `status='done', chunk_count=0`. + +Confirmed in server logs: +``` +NameError: name 'aiohttp' is not defined + at upload_file.dspy line 160: async with aiohttp.ClientSession(...) +``` + +**StreamHttpClient API** (pre-loaded, no import needed): + +```python +# Simple JSON POST — returns bytes, parse with json.loads() +client = StreamHttpClient() +resp = await client.request('POST', url, json={"key": "value"}) +result = json.loads(resp) + +# File upload (multipart) +resp = await client.request('POST', url, files={'field': (filename, file_bytes)}) + +# No timeout parameter — StreamHttpClient handles retry internally +# No context manager — use as a plain object, each call is independent +``` + +**Full example** — face detection in DSPY: +```python +try: + client = StreamHttpClient() + resp = await client.request('POST', 'https://media.opencomputing.net/face/api/detect', + json={"images": [img_b64]}) + fd = json.loads(resp) + results = fd.get("results", []) + if results and isinstance(results[0], dict): + face_count = len(results[0].get("faces", results[0].get("detections", []))) +except: pass +``` + +**Voiceprint with file upload**: +```python +try: + client = StreamHttpClient() + resp = await client.request('POST', 'https://media.opencomputing.net/voiceprint/extract/submit', + files={'file': (file_name, file_data)}) + vd = json.loads(resp) + voice_speakers = vd.get('speakers', 1) if vd.get('status') == 'SUCCEEDED' else 0 +except: pass +``` + +**Verification**: After any upload DSPY change, upload a test `.txt` file and query `document_chunks` to confirm it gained a row. Never rely on `status='done'` alone — check `chunk_count > 0`. + +**Anti-pattern — bare `except: pass` buries NameErrors**: +```python +# ❌ DANGEROUS — NameError silently swallowed +try: + client = aiohttp.ClientSession(...) # NameError +except: pass # all processing silently skipped + +# ✅ CORRECT — StreamHttpClient is pre-loaded, no NameError possible +try: + client = StreamHttpClient() + resp = await client.request('POST', url, json=data) +except: pass +``` + +## Bootstrap / Re-sync Local Repos + +When the local environment has NO ragserver or rag repos (first setup, or after cleanup), bootstrap them from remote and sync any test-server-only files: + +```bash +# 1. Create work dir and clone both repos from remote +mkdir -p /d/ymq/rag +cd /d/ymq/rag +git clone git@git.opencomputing.cn:yumoqing/ragserver.git ragserver +git clone git@git.opencomputing.cn:yumoqing/rag.git rag + +# 2. Audit test server for uncommitted/untracked files +ssh rag@rag.opencomputing.cn "cd /d/rag/ragserver && git status --short" +ssh rag@rag.opencomputing.cn "cd /d/rag/ragserver/pkgs/rag && git status --short" + +# 3. scp any server-only files to local +scp rag@rag.opencomputing.cn:/d/rag/ragserver/scripts/init_rbac_v2.py /d/ymq/rag/ragserver/scripts/ +scp rag@rag.opencomputing.cn:/d/rag/ragserver/pkgs/rag/wwwroot/.../new_file.dspy /d/ymq/rag/rag/wwwroot/.../ + +# 4. Stage, commit, push from local +cd /d/ymq/rag/ragserver # or rag +git add ... ; git commit -m "sync: ..." ; git push origin main + +# 5. Server pull (with pre-flight cleanup — see below) +``` + +**Clean up old duplicate repos** after bootstrap: remove any stale local clones under `/d/ymq/rag*`, `/d/ymq/rag-review`, `/d/ymq/repos/rag` so there is exactly ONE canonical local copy at `/d/ymq/rag/`. + +## Deployment Pitfalls + +### 🔴 Untracked files on server block git pull + +When a file exists on the server as untracked and the remote now tracks it (from a recent commit), `git pull` fails: +``` +error: The following untracked working tree files would be overwritten by merge: + scripts/init_rbac_v2.py +Please move or remove them before you merge. +Aborting +``` + +**Fix**: delete the untracked file on the server before pulling: +```bash +ssh rag@rag.opencomputing.cn "cd /d/rag/ragserver && rm -f scripts/init_rbac_v2.py && git pull" +``` + +This happens during the bootstrap sync workflow: the server had the file first (untracked), then you commit it locally and push. The server's copy must be removed so git can place the tracked version. + +### Config overwritten by git pull + +`conf/config.json` has local changes on the test server (encrypted password, driver, pool params removed). After `git pull`, these are reset. Also, **git pull silently fails** when local config.json has uncommitted changes — the server runs old code. + +**Pattern**: always use `git checkout conf/config.json && git pull` before deploy, then re-apply config fixes: + +```bash +ssh rag@rag.pd4e.com "cd ~/ragserver && git checkout conf/config.json && git pull && source py3/bin/activate && python3 -c \" +import json;c=json.load(open('conf/config.json')) +c['databases']['rag']['driver']='mysql' +c['databases']['rag']['kwargs']['password']='cybEz86hASifn+iwFHirOQ==' +c['databases']['rag']['kwargs'].pop('minsize',None) +c['databases']['rag']['kwargs'].pop('maxsize',None) +json.dump(c,open('conf/config.json','w'),indent=4,ensure_ascii=False) +\" && ps aux|grep 'app/ragserver'|grep -v grep|awk '{print \$2}'|xargs kill -9 2>/dev/null;sleep 1&&export PYTHONPATH=\$PWD:\$PWD/pkgs/rag-pipeline&&source py3/bin/activate&&nohup py3/bin/python app/ragserver.py>logs/startup.log 2>&1&sleep 3&&curl -so /dev/null -w '%{http_code}' http://localhost:9181/&&echo' ok'" +``` + +### RBAC cache clearing after manual SQL permission insert + +When permissions are added directly via SQL (not `init_rbac.py`), the in-memory cache may NOT refresh even after process restart. The `load_roleperms` method has a guard: if the DB returns 0 records (can happen during restart race), it keeps the previous cache with a debug message `'got 0 records, keeping previous cache'`. Full cache clearing: + +```bash +cd /d/rag/ragserver +bash stop.sh +pkill -9 -f ragserver.py # ensure no orphans +sleep 2 +find pkgs -name __pycache__ -exec rm -rf {} + 2>/dev/null +redis-cli FLUSHDB # clear session + RBAC caches +bash start.sh +``` + +The permission cache TTL is 10 minutes. Flushing Redis ensures both the RBAC `rp_caches` and user session state start fresh. + +### Table column naming: users.orgid vs knowledge_bases.org_id + +`users` table uses `orgid` (no underscore). `knowledge_bases` uses `org_id` (with underscore). `get_userorgid()` returns the user's `orgid`. Admin user's orgid is NOT "0" — on ragserver it's `4772b9b7031b4676`. When filtering by org in DSPY queries, use the correct column name for each table. + +```bash +ln -sf ../pkgs/rbac/wwwroot wwwroot/rbac +``` + +Without this, `/rbac/user/login.ui` returns 500 (file not found). + +### Session persistence: add session_max_time + session_issue_time + +Without these in `conf/config.json`, Redis sessions are never written and login is lost on restart: + +```json +"session_max_time": 3000, +"session_issue_time": 2500, +``` + +### 🔴 CRITICAL: site-packages copy stale after source edits — always reinstall + +When editing ANY Python module in `pkgs//` (rag, llmage, sqlor, etc.), the SERVER imports from `py3/lib/python3.10/site-packages//`, NEVER from `pkgs/`. Source edits have ZERO effect until reinstalled. **After ANY source change:** `cd /d/apitest/sage && ./py3/bin/pip install --upgrade /d/apitest/sage/pkgs/` then restart Sage. This applies to ALL modules — `pip install -e .` symlinks can break silently; `--upgrade` is more reliable. + +```bash +# Check if site-packages has the new function +grep 'extract_voiceprint' /d/rag/ragserver/py3/lib/python3.10/site-packages/rag/pipeline.py +# If NOT found, copy the source +cp /d/rag/ragserver/pkgs/rag/rag/pipeline.py /d/rag/ragserver/py3/lib/python3.10/site-packages/rag/pipeline.py +``` + +**Also try `pip install -e .` for new modules:** +```bash +cd /d/rag/ragserver/pkgs/rag +/d/rag/ragserver/py3/bin/pip install -e . +``` +Then restart. Without this, `from rag.pipeline import ...` → `ModuleNotFoundError` for brand-new `.py` files. + +**Symptom**: `ImportError: cannot import name 'X' from 'rag.pipeline'` — despite the function clearly existing in the source file. The site-packages copy is stale. + +### 🔴 ServerEnv registration — DOES NOT WORK for DSPY context (proven in production) + +**Full reference**: `references/dspy-context-imports.md` — definitive analysis of what works and what doesn't in DSPY exec context. + +Registering functions on `ServerEnv` (`env.func = func` in `init_rag_module()`) does **NOT** make them available in DSPY files. The DSPY exec context inherits pre-loaded globals (`json`, `uuid`, `DBPools`, `get_sor_context`, `params_kw`, `request`) from `globalEnv.py`, but ServerEnv attributes set by `init_rag_module()` do NOT propagate to `request._run_ns` in DSPY. + +**Verified failure pattern** (init_rag_module sets the function, server starts without error): +```python +def init_rag_module(): + env = ServerEnv() + from .pipeline import process_upload + env.process_upload = process_upload # set on ServerEnv singleton + rf = RegisterFunction() + ... +``` + +**DSPY calls it — ALWAYS returns NoneType:** +```python +env = request._run_ns +result = await env.process_upload(env, ...) # TypeError: 'NoneType' object is not callable +``` + +**The ONLY reliable zero-import pattern**: inline ALL logic directly in the DSPY file. No external function calls, no ServerEnv dependencies, no `from rag.pipeline import`. Use only ahserver pre-loaded globals (`json`, `uuid`, `DBPools`, `get_sor_context`, `request.read()`, `params_kw`). + +**For text extraction from documents** (PDF/DOCX/PPTX/XLSX): `PyPDF2`, `python-docx`, `python-pptx`, `openpyxl` ARE importable from DSPY — they're installed packages in the venv, not custom module code. Import them inline at point of use (not via separate pipeline.py). + +**Confirmed working pattern** (deployed to ragserver, commit `1cb82b9`): + +```python +# Place BEFORE the plain-text branch — office docs get priority +if ext_l == '.pdf' and not text: + import io; from PyPDF2 import PdfReader + reader = PdfReader(io.BytesIO(file_data)) + text = '\n'.join(p.extract_text() or '' for p in reader.pages) +elif ext_l == '.docx' and not text: + import io; from docx import Document + doc = Document(io.BytesIO(file_data)) + text = '\n'.join(p.text for p in doc.paragraphs) +elif ext_l == '.pptx' and not text: + import io; from pptx import Presentation + prs = Presentation(io.BytesIO(file_data)) + parts = [] + for slide in prs.slides: + for shape in slide.shapes: + if hasattr(shape, 'text') and shape.text: + parts.append(shape.text) + text = '\n'.join(parts) +elif ext_l == '.xlsx' and not text: + import io; from openpyxl import load_workbook + wb = load_workbook(io.BytesIO(file_data), data_only=True) + parts = [] + for sheet in wb.worksheets: + for row in sheet.iter_rows(values_only=True): + parts.append('\t'.join(str(c or '') for c in row)) + text = '\n'.join(parts) +``` + +**!!! NO `filetxt` package needed** — the inline pattern above is self-contained and uses only the four venv-installed packages. The `filetxt` module (git.opencomputing.cn:yumoqing/filetxt.git) exists as a reference but has heavy deps (spacy, langchain_community, mobi, ebooklib) and is NOT used in DSPY. + +**base64 in DSPY**: `import base64` works in DSPY files now that `ahserver/globalEnv.py` exposes `g.base64 = base64` (commit `93152f2`). DSPY files still need `import base64` at the top — it is NOT a pre-loaded global. Without the import, `NameError: name 'base64' is not defined` will still occur. + +**For RAG ingest**: The `_rag_ingest_async` function from init.py is NOT available in DSPY. Inline the UAPI calls directly in the DSPY using `StreamHttpClient` (pre-loaded global — no import needed). See `references/streamhttpclient-api.md` for patterns. + +**Why ServerEnv fails**: The DSPY exec wraps code in `async def myfunc(request, **ns):` with `exec(txt, lenv, lenv)`. `lenv` is populated from globalEnv pre-loaded globals, NOT from ServerEnv's dynamic attributes. The `run_ns.update(ServerEnv._ns_)` merge in baseProcessor happens at request time but the DSPY func's closure captures `lenv` at exec time — custom ServerEnv attributes aren't in scope. + +### Never delete __pycache__ blindly — the .pyc may be the only working version + +When debugging a Python service, do NOT run `rm -rf __pycache__` before checking if the `.py` source is complete. The `.pyc` file may contain compiled code from a PREVIOUS version that had server startup logic, while the current `.py` file may be a stripped-down version (e.g., just model functions without `__main__`). Deleting `__pycache__` can permanently break a working service. + +When a new `.py` file is added to the module package (e.g., `rag/pipeline.py`), the server can't find it until the package is reinstalled: +```bash +cd /d/rag/ragserver/pkgs/rag +/d/rag/ragserver/py3/bin/pip install -e . +``` +Then restart the server. Without this step, `from rag.pipeline import ...` → `ModuleNotFoundError`. + +Missing `.tmpl` processor causes `'NoneType' object has no attribute 'be_call'` on bricks header template load. + +### 🔴 DSPY HTTP calls: use StreamHttpClient, NOT aiohttp + +`StreamHttpClient` IS a pre-loaded DSPY global (`g.StreamHttpClient = StreamHttpClient`). `aiohttp` is NOT available in DSPY context. Using `aiohttp.ClientSession(...)` without import raises `NameError`. Critical danger: if wrapped in bare `except: pass`, the NameError is silently swallowed — the file upload succeeds but EVERY processing step fails silently: no embedding, no face detection, no voiceprint, no VDB upsert. Result: `status='done', chunk_count=0`. + +**StreamHttpClient API** (no import needed — pre-loaded global): +```python +# JSON POST — returns bytes, parse with json.loads() +client = StreamHttpClient() +resp = await client.request('POST', url, json={"key": "value"}) +result = json.loads(resp) + +# File upload (multipart) +resp = await client.request('POST', url, files={'field': (filename, file_bytes)}) +``` + +**Always wrap in try/except** — if the remote service is unreachable, the DSPY crashes 500 otherwise: +```python +try: + client = StreamHttpClient() + resp = await client.request('POST', url, json=payload) + result = json.loads(resp) +except: + result = {} +``` + +Pitfall: `StreamHttpClient.request()` returns raw bytes — no `.status`, no `.json()`. Must `json.loads(resp)`. + +### 🔴 buildUrlwidgetHandler: 4-parameter signature, not 3 + +When calling `bricks.buildUrlwidgetHandler()` from a script action, the function takes 4 parameters: +```javascript +// bricks.js line 465 +bricks.buildUrlwidgetHandler = function(w, target, rtdata, desc){ + var options = objcopy(desc.options||{}); // ❌ TypeError if desc is undefined +``` + +- `w` — source widget (`this` in script without `target`) +- `target` — from `getWidgetById` +- `rtdata` — data object (`{}` if no datawidget) +- `desc` — `{options: {url: ...}, mode: 'replace'}` (NOTE: `mode` goes INSIDE desc) + +❌ **Wrong (3 args, desc=undefined → TypeError):** +```javascript +buildUrlwidgetHandler({options:{url:u}}, target, 'replace') +``` + +✅ **Correct (4 args):** +```javascript +var t = bricks.getWidgetById('results', bricks.app.root); +if (t) bricks.buildUrlwidgetHandler(this, t, {}, {options:{url:u}, mode:'replace'}); +``` + +Full reference: `references/bricks-patterns-learned.md` + +### 🔴 search_result.dspy: MUST have try/except around HTTP calls + +The shell's `wwwroot/rag` **MUST** be a symlink targeting `../pkgs/rag/wwwroot`. This is the SINGLE MOST COMMON cause of mysterious 500 errors after deployments, git operations, or server restarts. **Check this FIRST before debugging anything else.** + +If `wwwroot/rag` is a regular directory (not a symlink), ALL `/rag/knowledge_bases_list/` paths fail with **500 "invalid path"** — every DSPY, every .ui, every page. Other routes (`/api/status`, `/`) work fine, making it look like a routing issue when it's not. + +**Symptom**: every page under `/rag/` returns 500, but other routes (`/api/status`, `/`) work fine. + +**Diagnose**: +```bash +ls -la wwwroot/rag +# Should show: rag -> ../pkgs/rag/wwwroot +# If it's a directory (drwxr-xr-x), the symlink is broken +ls wwwroot/rag/knowledge_bases_list/ +# Should list all .ui/.dspy files; if empty or missing, symlink is broken +``` + +**Fix**: +```bash +rm -rf wwwroot/rag +ln -sf ../pkgs/rag/wwwroot wwwroot/rag +``` +No server restart needed — file resolution happens on each request. + +### POST request params_kw: body consumed by getPostData before handler + +For POST requests, ahserver's `getArgs()` calls `getPostData(request)` instead of reading `request.query` directly. `getPostData` tries three body-reading strategies in order: + +1. `request.multipart()` — fails for raw binary (not multipart form data) +2. `request.post()` — returns empty dict for raw binary (not form-urlencoded) +3. `request.read()` — **consumes the request body** + +The query string IS merged into `params_kw` via `multiDict2Dict(request.query)` before the body read, so GET-style query params like `?folder=xxx` DO appear in `params_kw` even for POST requests. + +**Consequence for upload handlers**: when the handler later calls `await request.read()`, the body may already be consumed (depending on aiohttp's internal buffering). If the file data is still received (some aiohttp versions buffer), the handler works but the body was read twice. + +**Debugging**: when `params_kw.get("some_key")` returns empty but the Network tab shows the key in the URL, add the suspect key to the handler's response JSON to verify what was received: +```python +return json.dumps({ + "status": "SUCCEEDED", + "folder_id_received": folder_id, # <-- debug field + ... +}) +``` +Then check the Network response panel to see the actual value. If empty but the request URL had it, the issue is upstream in getPostData or DictObject conversion — not in your handler code. + +**Location**: `pkgs/ahserver/ahserver/processorResource.py`, method `getPostData` + `getArgs`. See lines ~265-295. + +### DSPY file placement vs URL path — MUST match wwwroot layout + +The `wwwroot/rag` symlink (`-> ../pkgs/rag/wwwroot`) means module pages at `pkgs/rag/wwwroot/knowledge_bases_list/` are accessed via `/rag/knowledge_bases_list/` URL, **NOT** `/knowledge_bases_list/`. When adding new dspy files: +- Place file: `pkgs/rag/wwwroot/knowledge_bases_list/new.dspy` +- URL path: `/rag/knowledge_bases_list/new.dspy` +- RBAC permission: `/rag/knowledge_bases_list/new.dspy` +- In `.ui` / `.dspy` files: `{{entire_url('/rag/knowledge_bases_list/new.dspy')}}` + +**Pitfall**: Using `/knowledge_bases_list/new.dspy` (without `/rag/` prefix) → 500 "invalid path" because no such file in `wwwroot/knowledge_bases_list/`. + +### RBAC permission path must match exact URL — verify with test + +After adding permissions, always verify with a real curl test (with session cookie if logined): +```bash +# 1. Login to get cookie +curl -s -c /tmp/cookies.txt -X POST "http://localhost:9181/rbac/user/up_login.dspy" \ + -d "username=admin&password=admin123" +# 2. Test the dspy with the cookie +curl -s -b /tmp/cookies.txt "http://localhost:9181/rag/knowledge_bases_list/storage_card.dspy?_webbricks_=1" +``` +401 = auth missing, 403 = permission wrong (path mismatch), 500 = dspy code bug, 200 = working. + +### Always curl-test new DSPY endpoints after deployment + +**User mandate**: Every new dspy MUST be tested via curl before declaring done. Do NOT just deploy and assume it works. Test from the server itself (localhost) with a session cookie. Check the server log (`logs/ragserver.log`) for exceptions. 500 errors are invisible in the browser (bricks swallows them) but logged server-side. + +### New DSPY files MUST be added to init_rbac.py — or get 403 + +Every new `.dspy` file under a module's `wwwroot/` directory needs an explicit path entry in `scripts/init_rbac.py` (PUBLIC or LOGINED list). The directory-level entry (`/rag/knowledge_bases_list/`) does NOT auto-cover files within — each file needs its own permission row. + +**Checklist when adding a new `.dspy`:** +1. Add the path to `init_rbac.py` PUBLIC or LOGINED list +2. Run `cd /d/rag/ragserver && py3/bin/python3 scripts/init_rbac.py` +3. **Verify the rolepermission association was created** — the script uses `INSERT IGNORE`, and permission ID truncation (see below) causes the rolepermission insert to silently fail +4. Restart the server + +**Verification query:** +```bash +mysql -h db -u test -p'test123' rag -e " +SELECT p.id, p.path, rp.roleid +FROM permission p +LEFT JOIN rolepermission rp ON rp.permid = p.id +WHERE p.path LIKE '%your_new_file%' +" +``` +If `rp.roleid` is NULL, the rolepermission wasn't created — insert it manually. + +### Permission ID truncation (varchar 32) causes silent rolepermission failures + +`init_rbac.py` generates permission IDs as `perm_`. For long paths, this exceeds the `permission.id` column's varchar(32) limit. MySQL silently truncates the ID on insert, but the `rolepermission` insert uses the full (un-truncated) `rp_perm_` as its ID and `perm_` as `permid`. Since the `permid` references the truncated permission ID, the `INSERT IGNORE` fails silently (FK constraint or similar). + +**Example**: Path `/rag/knowledge_bases_list/file_list.dspy` generates +- `pid = perm__rag_knowledge_bases_list_file_list.dspy` (45 chars) +- Stored as `perm__rag_knowledge_bases_list_f` (truncated to 32) +- `rp_perm__rag_knowledge_bases_list_file_list.dspy` tries to link to truncated `perm__rag_knowledge_bases_list_f` — mismatch causes silent skip + +**Fix**: After running `init_rbac.py`, always verify with the query above. If missing, insert manually: +```sql +INSERT IGNORE INTO rolepermission (id, roleid, permid) +VALUES ('rp_', 'any', ''); +``` +Where `` is the actual stored permission ID (first 32 chars). + +### User passwords: RC4, not AES + +RBAC login uses `rc4.password(s, key=k)`. DB connection uses `aes_encode_b64`. They are different systems — don't mix them up. + +### Hardcoded storage display — "已用 0MB" always shows zero + +The KB list page `knowledge_bases_list/index.ui` hardcoded `{"text": "已用 0MB / 总计 100MB"}` — never queries DB. Fix: replace with `urlwidget` loading a DSPY. + +**storage_card.dspy**: +```python +ns = params_kw.copy() +env = request._run_ns + +async with get_sor_context(env, 'rag') as sor: + rec = await sor.sqlExe("SELECT COALESCE(SUM(total_size),0) used FROM knowledge_bases", {}) + used_bytes = int(rec[0].used) if rec else 0 + +used_mb = round(used_bytes / 1048576, 1) +pct = min(round(used_bytes / 104857600 * 100), 100) if used_bytes > 0 else 0 + +return { + "widgettype": "VBox", + "options": {"width": "100%", "bgcolor": "#f8f9fa", "padding": "16px", "css": "card", "spacing": "4px"}, + "subwidgets": [ + {"widgettype": "HBox", "options": {"width": "100%", "justifyContent": "space-between"}, "subwidgets": [ + {"widgettype": "Text", "options": {"text": "💾 存储容量", "cfontsize": 14, "fontWeight": "bold", "color": "#333"}}, + {"widgettype": "Text", "options": {"text": "已用 " + str(used_mb) + "MB / 总计 100MB", "cfontsize": 12, "color": "#888"}} + ]}, + {"widgettype": "VBox", "options": {"width": "100%", "cheight": 0.5, "bgcolor": "#e0e0e0"}, "subwidgets": [ + {"widgettype": "VBox", "options": {"cwidth": pct, "cheight": 0.5, "bgcolor": "#4a90d9"}} + ]} + ] +} +``` + +**index.ui patch**: Replace the hardcoded VBox card with: +```json +{"widgettype": "urlwidget", "options": {"url": "{{entire_url('/rag/knowledge_bases_list/storage_card.dspy')}}"} +``` + +### detail.ui architecture — tree + file_list.dspy with node_selected + +The knowledge base detail page uses a split layout: +- **Left**: Tree widget (`dir_tree`) with editable CRUD for folder management +- **Right**: urlwidget (`file_list_panel`) loading `file_list.dspy` for the current folder +- **Bind**: tree's `node_selected` event reloads `file_list.dspy` with the selected node's `{id, label}` as query params + +```json +{ + "widgettype": "VBox", + "id": "detail_content", + "options": {"css": "filler", "padding": "16px", "spacing": "0"}, + "subwidgets": [ + { + "widgettype": "urlwidget", + "id": "file_list_panel", + "options": { + "url": "{{entire_url('./file_list.dspy')}}?kb_id={{params_kw.kb_id}}&id=__root__" + } + } + ], + "binds": [ + { + "wid": "dir_tree", + "event": "node_selected", + "actiontype": "urlwidget", + "target": "file_list_panel", + "mode": "replace", + "options": { + "url": "{{entire_url('./file_list.dspy')}}?kb_id={{params_kw.kb_id}}" + } + } + ] +} +``` + +`file_list.dspy` reads `kb_id`, `id` (folder_id), `label` from params_kw, queries documents for the given folder, and renders upload UI + file list. + +**Pitfall**: detail.ui was modified on the server but overwritten by `git pull` because changes weren't committed. Always `git add && git commit && git push` after modifying .ui/.dspy files on the server. + +### RBAC permission path MUST include `/rag/` prefix (post fbba95e) + +The `fbba95e` commit added `/rag/` prefix to all module URLs. Permissions created before this commit (without prefix) won't match. When adding new permissions: +- Path in DB: `/rag/knowledge_bases_list/.dspy` (WITH `/rag/`) +- In `.ui` files: `{{entire_url('/rag/knowledge_bases_list/.dspy')}}` + +### DSPY URL encoding — Chinese chars in query params break ahserver + +ahserver rejects unencoded non-ASCII characters in URL query strings. `entire_url()` passes values literally — it does NOT percent-encode. Passing Chinese like `&label=根目录` causes `Invalid char in url query` error. + +**Fix**: Keep all DSPY URL query params ASCII-only. Derive Chinese labels internally: + +```python +# ❌ DON'T — Chinese in URL query +url = entire_url('./page.dspy') + '&label=根目录&id=__root__' + +# ✅ DO — ASCII-only, derive label inside DSPY +url = entire_url('./page.dspy') + '&id=__root__' +# In page.dspy: ns = params_kw.copy(); label = '根目录' if ns.get('id') == '__root__' else ns.get('label', '') +``` + +**Always curl-test after adding permissions**: +```bash +# Login, then verify each endpoint +curl -s -b "$COOKIE" "http://localhost:9181/rag/knowledge_bases_list/file.dspy?_webbricks_=1" | head -c 200 +``` +401 = auth, 403 = permission path mismatch, 500 = dspy bug, 200 = OK. + +### RBAC dual-path hell — cleanup for /knowledge_bases_list/ requires login + +After the `/rag/` prefix was added, knowledge_bases_list pages should ALL require login (`logined` role), not public (`any` role). Use `scripts/init_rbac_v2.py` which auto-scans wwwroot and assigns `logined` to all business pages. + +This ensures unauthenticated users cannot browse knowledge bases. Public endpoints (login, register, status, bricks) remain in `any`. + +### Jinja2 backslash template error — \** inside {{}} breaks .ui rendering + +`.ui` files processed by the `bui` processor go through Jinja2 template rendering. Backslash-escaped quotes (`\"`) inside `{{}}` expressions cause `TemplateSyntaxError: unexpected char '\\'`. This happens when HTML/JS snippets with inline event handlers are embedded in JSON strings. + +**Fix**: Remove backslashes before quotes INSIDE `{{...}}` blocks: +```sql +-- Add logined for all KB paths +INSERT IGNORE INTO rolepermission (id, roleid, permid) +SELECT CONCAT('rp_', p.id), 'logined', p.id +FROM permission p +WHERE p.path LIKE '%knowledge_bases_list%' +AND NOT EXISTS (SELECT 1 FROM rolepermission rp WHERE rp.permid = p.id AND rp.roleid = 'logined'); + +-- Remove any public (any) access to KB pages +DELETE rp FROM rolepermission rp JOIN permission p ON rp.permid = p.id +WHERE p.path LIKE '%knowledge_bases_list%' AND rp.roleid = 'any'; + +-- Remove ghost permissions (files that no longer exist) +DELETE FROM rolepermission WHERE permid IN ( + SELECT id FROM permission WHERE path NOT LIKE '/api/%' AND path NOT LIKE '/bricks/%' + AND path NOT LIKE '/i18n/%' AND path NOT LIKE '/_%' +); +``` + +Then flush Redis + restart: +```bash +redis-cli FLUSHDB +cd /d/rag/ragserver && bash stop.sh && pkill -9 -f ragserver.py; sleep 2; bash start.sh +``` + +### Ghost permissions — deleted files leave orphaned DB entries + +When a `.dspy` or `.ui` file is deleted from `wwwroot/`, its permission entries in the `permission` and `rolepermission` tables remain. These ghost permissions cause confusion: curl returns 500 (file not found) but RBAC passes (403 would mean permission issue). Always verify file existence when debugging 500s: +```bash +ls /d/rag/ragserver/pkgs/rag/wwwroot/knowledge_bases_list/get_kb_cards.dspy +# vs +mysql ... -e "SELECT * FROM permission WHERE path LIKE '%get_kb_cards%'" +``` +If the file doesn't exist but the permission does → ghost. Delete the permission. + +### Git pull overwrites uncommitted server changes — user mandate + +**User preference (explicit)**: "代码永远在本地改,提交远程后,测试服务器git pull这样不会出问题". Never manually edit files on the server. Edit locally → `git commit` → `git push` → server `git pull`. Manual server edits are overwritten by `git pull` and cannot be recovered because they were never tracked. +```bash +cd /d/rag/ragserver/pkgs/rag +git stash # save manual changes before pull +git pull +git stash pop # restore after +``` +Better: never manually edit on the server — edit locally, push, then pull on server. + +### Two-SSH-hop deployment pattern (when local can't reach target directly) + +When the local machine can't SSH directly to the rag server but can reach an intermediate: +```bash +scp /tmp/file apitest@120.48.168.15:/d/apitest/ +ssh apitest@120.48.168.15 "scp /d/apitest/file rag@rag.opencomputing.cn:/tmp/" +ssh apitest@120.48.168.15 "ssh rag@rag.opencomputing.cn 'command'" +``` + +### urlwidget relative URL pitfall — use absolute paths in scripts + +When a `.ui` page is loaded as a sub-widget via `urlwidget` within the shell, the browser's base URL is the shell page, NOT the sub-widget's URL. Relative URLs like `./upload_file.dspy` in XHR/script resolve to the ROOT, not to the sub-widget directory. **Always use absolute paths** in script XHR calls: `/rag/knowledge_bases_list/upload_file.dspy`. + +### Upload completion refresh — put refreshList AFTER if/else + +When uploading multiple files, the refresh call must be placed OUTSIDE the if/else block so it fires regardless of success/failure, after all uploads complete: + +```javascript +x.onload=function(){ + done++; + if(x.status===200){ st.innerText='上传成功 '+done+'/'+total } + else{ st.innerText='上传失败' } + if(done>=total)refreshList() // outside if/else — always runs after last upload +}; +``` + +### Voiceprint ahserver startup — auth bypass needed for dependency compatibility + +When starting voiceprint with ahserver (`python3 ah.py -p 9087`), the imported `ahserver` may use a different `aiohttp_auth` version than what's installed. The installed `aiohttp_auth.auth` module may lack the `setup()` method, and `auth.get_auth()` may throw `RuntimeError: auth_middleware not installed`. + +**Fix — two changes needed**: +1. In `ahserver/ahserver/auth_api.py`, add `return` at the start of `setupAuth()` (skip auth middleware setup) +2. In `ahserver/ahserver/processorResource.py`, change `self.user = await auth.get_auth(request)` to `self.user = None # auth disabled` + +**Dependencies needed** (install once): +```bash +pip install aiohttp_auth aiohttp_session aiohttp_cors aiohttp_middlewares openpyxl rsa qrcode asyncssh +``` + +**Start command**: +```bash +cd /share/ymq/run/voiceprint +PYTHONPATH=.:sqlor:appPublic:ahserver:longtasks/longtasks \ + nohup python3 ah.py -p 9087 > logs/voiceprint.log 2>&1 & +``` + +**Verify**: `curl http://localhost:9087/api/status` → should return JSON with `"ready": true` + +### Face API parameter format — `images` (array), NOT `image` + +The face-service accepts `{"images": ["base64_str", ...]}` — an array of base64 strings. Sending `{"image": "..."}` (singular, non-array) returns 500. +``` diff --git a/skills_library/all/rbac-permission-initialization-pattern/SKILL.md b/skills_library/all/rbac-permission-initialization-pattern/SKILL.md new file mode 100644 index 0000000..0736d48 --- /dev/null +++ b/skills_library/all/rbac-permission-initialization-pattern/SKILL.md @@ -0,0 +1,925 @@ +--- +name: rbac-permission-initialization-pattern +description: Pattern for initializing RBAC permissions in business modules that use wildcard expansion, role ID matching, and path registration +author: Hermes Agent +tags: [rbac, permissions, init, multi-tenant, sage] +--- + +# RBAC Permission Initialization Pattern + +## Overview +When deploying a business module with RBAC authentication, the permission initialization must handle several complex scenarios: path normalization, role wildcard expansion, CRUD file structure, and URL rewriting edge cases (WSS, index auto-match). + +**User preference**: Each business module owns its own `scripts/load_path.py` that registers its paths directly via Sage DB operations (mirroring Sage's `load_path.py` internals). This keeps permissions self-contained per module and runnable from any Sage environment. See `references/per-module-load-path.md` for the Python template pattern. The Sage-level `load_path.py` is the canonical declarative source of truth, but for per-module workflows, the module's own script is preferred. + +## Role-Based Permission Analysis Methodology + +Before writing permission scripts, analyze each role's responsibilities and classify paths into tiers: + +**Step 1: Document role职责** +``` +owner.superuser — 系统级: 机构类型/角色/权限管理, 添加业主管理员 +*.admin — 机构级: 添加本机构人员, 分配人员角色 +reseller.operator — 运营: 产品管理/供应商合同/定价/统一折扣/营销 +reseller.sale — 销售: 客户管理/客户特殊折扣 +reseller.accountant — 财务: 线下充值/对账结算 +reseller.maintainer — 运维维护 +customer.customer — 终端客户用户 +logined — 所有已登录用户 +``` + +**Step 2: Analyze module business nature** +- Ask: Is this a domain-specific business module (CRM, accounting) or a general tool service (AI agent, reasoning)? +- General tool services → broader access (all logined users can use) +- Domain-specific modules → restricted to relevant roles + +**Step 3: Classify into permission tiers** +| Tier | Role set | Path types | +|------|----------|-----------| +| **Public** | `any` | 登录/注册/认证页面、静态资源(img/css) | +| **Logined** | all 登录角色 | 用户自助服务(个人信息、API Key)、数据查看(列表+get)、用户自己的CRUD(通过user_id隔离) | +| **Admin** | superuser + *.admin | 系统/机构配置管理、用户管理、机构管理 | +| **Superuser** | `owner.superuser` only | 全局元数据(角色/权限/机构类型)、高危操作(技能部署) | + +**Step 4: Register CRUD paths comprehensively** +- JSON CRUD `alias` → directory `{alias}/` with `index.ui`, `get_*.dspy`, `add_*.dspy`, `update_*.dspy`, `delete_*.dspy` +- Custom `api/` directory may also contain CRUD endpoints — register both +- Every CRUD directory needs TWO paths (see Pitfall 2) + +## Key Principles + +### 1. Wildcard Expansion (Filesystem Scanning) +**Problem**: `rbac.check_roles_path()` does exact matching. Permission definitions must be registered as concrete URLs in the DB. + +**CRITICAL**: Scan ALL file types — `.ui`, `.dspy`, `.js`, `.css`. RBAC protects all static resources. + +**IMPORTANT: ahserver auto-serves `.css` and `.js` files** from module `wwwroot/` directories — they are injected into HTML responses without explicit ``/`