--- name: dspy-file-implementation-spec description: Standardized specification for implementing .dspy files in ahserver applications with proper return format and module integration author: Hermes Agent tags: [ahserver, dspy, backend, web-development, python] --- # .dspy File Implementation Specification ## Overview .dspy files are controlled Python scripts executed by the ahserver web framework to provide dynamic API endpoints. They must follow strict conventions to ensure security, performance, and compatibility with the framework's architecture. ## Core Principles 1. **No Import Statements** **Never use import statements** in .dspy files. The ahserver framework: - Automatically provides access to functions exported by your application module through `load_{modulename}()` - Has already pre-loaded common Python modules (datetime, json, os, sys, etc.) into the global context **❌ Incorrect:** ```python import json import datetime from datetime import date, timedelta from myapp.init import get_all_records ``` **✅ Correct — use pre-loaded modules directly:** ```python # datetime is pre-loaded as the full module — access via datetime.date, datetime.datetime, datetime.timedelta today = datetime.date.today().isoformat() now = datetime.datetime.now() five_min_ago = (now - datetime.timedelta(minutes=5)).strftime('%Y-%m-%d %H:%M:%S') # json is pre-loaded result = json.dumps({'key': 'value'}) # Directly use functions provided by load_app_module() records = get_all_records() ``` **⚠️ Pitfall**: `from datetime import date` looks innocent but WILL cause the .dspy file to fail with an import error. Use `datetime.date.today()` instead. ### 2. Use Return, Not Print **Always use `return` to send data back to the client**, never use `print()`. The ahserver framework handles JSON serialization automatically. **❌ Incorrect:** ```python result = {"data": records} print(json.dumps(result)) ``` **✅ Correct:** ```python return records ``` ### 3. ID Generation: `uuid()` in .dspy/.ui, `getID()` in .py **CRITICAL**: Both `uuid()` and `getID()` are available in `.dspy` context: ```python # Both work in .dspy context — use uuid() for new IDs (shorter, simpler) new_id = uuid() # getID() is also pre-loaded in .dspy context (verified: llmage dspy files use it without import) new_id = getID() ``` In `.py` files (e.g., `init.py`, `utils.py`), you must import: `from appPublic.uniqueID import getID`. ### 4. Proper Error Handling Handle exceptions gracefully and return appropriate data structures based on component requirements. **For array-returning endpoints (e.g., code components):** ```python try: records = get_all_records() result = [] for record in records: result.append({ "value": str(record.get('id')), "text": record.get('name', f"Record {record.get('id')}") }) return result except Exception as e: return [] # Return empty array on error ``` **For object-returning endpoints:** ```python try: record = get_record_by_id(id) return record except Exception as e: return {"error": str(e)} ``` ## Common Use Cases ### 1. Code Component Data Endpoints Code components require specific `{value, text}` array format: **File:** `/wwwroot/entity_name/list/index.dspy` ```python # Get entity list for code dropdown # This .dspy file uses functions released by load_app_module() try: # Use the function provided by your module records = get_all_records() # Format for code component (value, text pairs) result = [] for record in records: result.append({ "value": str(record.get('id')), "text": record.get('name', f"Record {record.get('id')}") }) # Return array directly for code component return result except Exception as e: # On error or no data, return empty array return [] ``` ### 2. Single Record Endpoints For retrieving individual records: **File:** `/wwwroot/entity_name/get/index.dspy` ```python # Get single entity record # Access query parameters via params_kw dictionary try: record_id = params_kw.get('id') if not record_id: return {"error": "ID parameter required"} record = get_record_by_id(record_id) return record except Exception as e: return {"error": str(e)} ``` ### 3. Action Endpoints For performing actions like testing connections: **File:** `/wwwroot/entity_name/test/index.dspy` ```python # Test entity connection or perform action try: entity_id = params_kw.get('id') if not entity_id: return {"status": "error", "message": "ID parameter required"} result = test_entity_connection(entity_id) return {"status": "success", "message": result} except Exception as e: return {"status": "error", "message": str(e)} ``` ### 4. Login Endpoint Pattern Login endpoints require special handling for password encoding and session creation: **File:** `/wwwroot/login.dspy` ```python #!/usr/bin/env python3 # -*- coding: utf-8 -*- """Login handler - uses server-env functions, no imports needed""" username = params_kw.get('username', '') password = params_kw.get('password', '') if not username: return json.dumps({'status': 'error', 'message': 'Username required'}, ensure_ascii=False) if not password: return json.dumps({'status': 'error', 'message': 'Password required'}, ensure_ascii=False) # Encode password for comparison with stored hash passwd = password_encode(password) # Use server-env registered check_user_password rzt = await check_user_password(request, username, passwd) if rzt: # Get user info from database dbname = get_module_dbname('rbac') async with DBPools().sqlorContext(dbname) as sor: users = await sor.sqlExe( "SELECT id, username, name, orgid FROM users WHERE username=${username}$", {'username': username} ) if users: user = users[0] # Create session using remember_user (available in .dspy context) await remember_user(user.id, user.username, getattr(user, 'orgid', '') or '') return json.dumps({ 'status': 'ok', 'message': 'Login successful', 'redirect': '/main/base.ui', 'userid': user.id, 'username': user.username }, ensure_ascii=False) # Failed login return json.dumps({'status': 'error', 'message': 'Invalid credentials'}, ensure_ascii=False) ``` **Key points for login .dspy:** - Use `password_encode()` to hash the submitted password before comparison - Use `check_user_password(request, username, encoded_password)` for RBAC authentication - Use `remember_user(userid, username, userorgid)` to create session (NOT `user_login()` - that requires explicit import which fails in .dspy) - Return a string via `json.dumps()`, never return `None` ## Security Considerations ### 1. Input Validation Always validate and sanitize input parameters from `params_kw`: ```python # Validate ID parameter record_id = params_kw.get('id') if not record_id or not str(record_id).isdigit(): return {"error": "Invalid ID parameter"} ``` ### 2. Avoid Sensitive Data Never return sensitive fields like passwords, API keys, or internal system data unless explicitly required and properly authorized. ### 3. Rate Limiting For production applications, implement rate limiting for expensive operations: ```python # Check rate limit (pseudo-code) if is_rate_limited(request_ip): return {"error": "Rate limit exceeded"} ``` ## Performance Guidelines ### 1. Efficient Data Retrieval Use appropriate database queries with proper filtering and pagination: ```python # Use efficient queries with limits records = get_records_with_limit(offset=0, limit=100) ``` ### 2. Caching Implement caching for frequently accessed, rarely changing data: ```python # Use application-level cache cache_key = f"records_list_{timestamp}" if cache_key in app_cache: return app_cache[cache_key] records = get_all_records() app_cache[cache_key] = records return records ``` ## DSPY Code Review Checklist A structured checklist for reviewing `.dspy` files — see `references/dspy-code-review-checklist.md` for detailed walkthroughs of each check with real-world bug examples (Decimal serialization crashes, missing `int()` on SUM aggregates, DRY violations, sibling-file inconsistency detection). ### Syntax & Security - [ ] **No imports** — module DSPY files must have zero import statements. All needed names (`json`, `datetime`, `get_sor_context`, `DBPools`, `params_kw`, `request`, `uuid`, `time`, `os`, `DictObject`, `FileStorage`, logging functions) are pre-loaded. - [ ] **No forbidden patterns** — no `eval()`, `exec()`, `__import__()`, `os.system()`, `subprocess`, `pickle.loads()`. - [ ] **Valid Python AST** — file passes `ast.parse()`. Quick check: `python3 -c "import ast; ast.parse(open('file.dspy').read()); print('OK')"`. **⚠️ .dspy files contain top-level `await`/`async with` which bare ast.parse rejects ("await outside async function")** — wrap first: `wrapped = 'async def __c__(params_kw, request, uid, org_id, json, DBPools, get_user, get_userorgid, get_module_dbname, getID, debug, sor, params_kw=None):\n' + '\n'.join(' ' + line if line.strip() else line for line in src.split('\n')); ast.parse(wrapped)` (add injected names the file uses to the wrapper signature). This wrapped check is **mandatory after patching triple-quoted prompt constants** — a stray `"""` silently closes the string and dumps the following prose as code; only ast.parse exposes it (caught live 2026-08 in cockpit_chat.dspy). - [ ] **All branches return** — every code path ends with an explicit `return`. Missing return → `return data type error, `. ### SQL & Database - [ ] **Parameterized queries** — uses `${param}$` syntax, never f-string interpolation or `%s` formatting in SQL strings. - [ ] **Decimal / SUM aggregate safety** — `SUM()` in MySQL returns `Decimal`. Must wrap with `int()` or pass `default=str` in `json.dumps()`. Check: `r.total_size or 0` should be `int(r.total_size or 0)`. This is the same class of bug as doc_count/chunk_count lacking `int()`. - [ ] **Cross-module access** — uses `get_sor_context(env, 'module')`, not `DBPools().sqlorContext(dbname)` for modules outside the current one. - [ ] **sqlExe return type awareness** — without `page`/`rows` in ns → list of row objects (use `r.field` attrs); with `page`/`rows` → `{'total': N, 'rows': [...]}` dict. - [ ] **Error handling** — at least a try/except around DB ops with a fallback return. ### Code Quality (KISS/DRY) - [ ] **Sibling file consistency** — compare against other `.dspy` files in the same directory. Inconsistent return format (raw dict vs `json.dumps()`), divergent helper signatures, or different API patterns are red flags. - [ ] **DRY — no duplicated helpers** — check for size formatters (`fmt_size`, `fmt`), date formatters, or SQL builders duplicated across files in the project. Three identical copies of the same function is a signal to extract. - [ ] **No hardcoded config values** — storage limits, API URLs, timeouts should come from config, not be embedded in code. - [ ] **f-string safety** — avoid f-strings in dict returns; `exec()` wrapping can misparse `}` braces. Use concatenation `'prefix: ' + str(var)` instead. - [ ] **No `print()`** — use `return` for output. `print()` writes to stdout that ahserver ignores, producing `NoneType` error. ### Return Format - [ ] **Consistent return style** — all DSPY files in a directory should use the same pattern: either raw dict `return {...}` or `json.dumps({...})`. - [ ] **DataViewer CRUD endpoints** — must return `Message` widget JSON, not raw data. - [ ] **Code component endpoints** — must return `[{value, text}]` array. - [ ] **JSON validity** — if the DSPY returns a hardcoded JSON-like dict, validate the resulting JSON serializes correctly (watch for `Decimal`, `datetime`, `bytes` types that `json.dumps` can't handle without `default=str`). ## Testing and Validation ### 1. Manual Testing Test .dspy endpoints directly by accessing their URLs in a browser: ``` http://localhost:8000/app-name/entity_name/list/ ``` ### 2. Data Format Validation Verify that returned data matches the expected format for the consuming component: - **Code components**: Array of `{value, text}` objects - **DataViewer**: Array of full record objects - **Forms**: Single record object or success/error object ### 3. Error Scenario Testing Test error scenarios like missing parameters, invalid IDs, and database failures. ## Integration with Bricks Framework ### 1. UI File References Reference .dspy endpoints in .ui files using standard URL format: ```json { "uitype": "code", "data_url": "/app-name/entity_name/list/" } ``` ### 2. Parameter Passing Pass parameters to .dspy endpoints using query strings: ```json { "data_url": "/app-name/entity_name/get/?id={{selectedRow.id}}" } ``` ## CRUD List API Pattern (sqlor-based) For DataGrid/CRUD widget data endpoints, use this standardized pattern: ```python # CRUD list API for DataViewer — no imports needed, json/DBPools are pre-loaded result = {'success': False, 'rows': [], 'total': 0} try: dbname = get_module_dbname('module_name') async with DBPools().sqlorContext(dbname) as sor: # Build WHERE clause dynamically where_clauses = [] where_ns = {} customer_id = params_kw.get('customer_id', '') status = params_kw.get('status', '') if customer_id: where_clauses.append("customer_id=${customer_id}$") where_ns['customer_id'] = customer_id if status: where_clauses.append("status=${status}$") where_ns['status'] = status where_sql = " AND ".join(where_clauses) where_prefix = " WHERE " if where_clauses else "" # Count query (no pagination needed) count_sql = "SELECT count(*) rcnt FROM table_name" + where_prefix + where_sql count_rows = await sor.sqlExe(count_sql, where_ns) total = 0 if count_rows and len(count_rows) > 0: r = count_rows[0] if hasattr(r, 'keys'): total = r.get('rcnt', 0) elif isinstance(r, dict): total = r.get('rcnt', 0) elif hasattr(r, 'rcnt'): total = r.rcnt if total > 0: # Pagination query ns = {'page': int(params_kw.get('page', 1)), 'rows': int(params_kw.get('rows', 20)), 'sort': params_kw.get('sort', 'id')} sql = "SELECT col1, col2, col3 FROM table_name" + where_prefix + where_sql # Merge ns and where_ns (avoid {**ns, **sql_ns} which fails) query_ns = dict(list(ns.items()) + list(where_ns.items())) rows = await sor.sqlExe(sql, query_ns) # sqlExe with page/rows returns {'total': N, 'rows': [...]} if isinstance(rows, dict): result['rows'] = rows.get('rows', []) result['total'] = rows.get('total', total) elif rows: result['rows'] = [dict(r) if hasattr(r, 'keys') else r for r in rows] result['total'] = total result['success'] = True except Exception as e: result['error'] = str(e) return json.dumps(result, ensure_ascii=False, default=str) ``` **Key points:** - Return format: `{'success': bool, 'rows': [...], 'total': int}` - Use `params_kw.get()` for pagination parameters - Use `${param}$` syntax for LIMIT/OFFSET in sqlExe - Convert rows to dicts: `[dict(r) for r in data]` - Use `default=str` in json.dumps for datetime handling - **CRITICAL**: All SELECT columns must match the actual database schema exactly. Always verify with `DESCRIBE table_name` before writing queries. ## Cross-Module Database Access Pattern When a .dspy file in one module needs to access tables belonging to another module: ### REQUIRED: `get_sor_context(request._run_ns, 'module')` — the ONLY correct pattern ```python # In .dspy files — request is auto-injected env = request._run_ns async with get_sor_context(env, "module_name") as sor: records = await sor.R('table_name', {'filter': 'value'}) ``` This is the **only** cross-db access pattern. It works because it delegates to the `module_dbname` config: in the Sage system, a module named "tenant" resolves to the `sage` database; in the pipeline-app, the same module resolves to the `pipeline` database. The module's owner configures this mapping per deployment. ### ❌ NEVER use hardcoded database names ```python # WRONG — hardcoded db name breaks cross-deployment portability async with db.sqlorContext("pipeline") as sor: ... ``` This is the single most common cross-module DSPY error. It works in one environment but fails in another (e.g., Sage queries "pipeline" DB which doesn't exist in its DBPools config). Always use `get_sor_context(env, "module_name")` instead. ### ❌ NEVER use `DBPools()` + `sqlorContext()` for cross-module access The `DBPools()` pattern is for accessing the **current** module's database. For cross-module access, use only `get_sor_context`. **Key points:** - **Never use `ServerEnv()` in .dspy files** — all server-env functions (`get_module_dbname`, `DBPools`, `getConfig`, `password_encode`, etc.) are already injected into the .dspy execution context via globals - **Never hardcode database names** in .dspy files — use `get_sor_context(env, "module_name")` to resolve via config - `get_sor_context(request._run_ns, 'module')` is the **required** pattern for cross-module DB access - If a cross-module function is registered via `load_{modulename}()` (like `create_user_apikey` from dapi), use it directly: `create_user_apikey(sor, dappid, user_id, user_orgid)` ## Batch Operations with $or Queries For batch lookups by ID list, use `$or` conditions in the sor.R filter: ```python # user_ids is a list of IDs to look up or_conditions = [{'id': uid} for uid in user_ids] query_ns = {'$or': or_conditions} users = await sor.R('users', query_ns) ``` **Key points:** - The `$or` operator is supported by sqlor's filter system - For large lists (>100 items), consider chunking to avoid query complexity limits - Always validate the ID list is non-empty before querying ## Safe Attribute Access on SQLor Row Objects SQLor returns row objects that may or may not support dict-style access. Use `getattr()` for safe attribute access: ```python user = users[0] user_id = getattr(user, 'id', '') username = getattr(user, 'username', '') user_orgid = getattr(user, 'orgid', '') or '' # Handle None -> '' ``` **Key points:** - `getattr(obj, 'attr', default)` is safer than `obj.attr` (avoids AttributeError) - Use `or ''` pattern for fields that may be None but need to be a string - For dict-like access: `getattr(user, 'orgid', '') or ''` handles both missing attribute and None value ## Server-Env Functions Available in .dspy Context The ahserver framework injects many functions into the .dspy execution context. **No import needed** - just use them directly: | Function | Description | |----------|-------------| | `password_encode(s)` | Hash a password using the app's configured key | | `password_decode(s)` | Decode a hashed password | | `remember_user(userid, username, userorgid)` | Set session user (login) | | `forget_user()` | Clear session user (logout) | | `get_user()` | Get current logged-in user ID | | `get_username()` | Get current user's display name | | `get_userorgid()` | Get current user's org ID | | `get_userinfo()` | Get full user info object | | `get_session()` | Get session object | | `session_getvalue(key)` | Read session value | | `session_setvalue(key, value)` | Write session value | | `get_module_dbname(modulename)` | Get DB name for a module | | `get_sor_context(env, modulename)` | Async context manager for cross-module DB access | | `DBPools()` | Get database connection pool | | `params_kw` | Dictionary of request parameters — query string + POST body (including `application/json`), merged into one dict. Nested JSON objects preserved as dict/list. **This is the ONLY way to access request data — there is NO `http_request` variable.** | | `request` | The ahserver Request object (auto-injected) | | `json` | json module (json.dumps, json.loads) | | `datetime` | datetime module (datetime.date, datetime.datetime, datetime.timedelta) | | `uuid` / `getID` | ID generation — both work. `uuid()` returns shorter IDs, `getID()` returns 22-char IDs | | `time` | time module | | `os` | os module (MAY be available — verify if needed; observed as imported in recover_usages.dspy for `os.path.isfile`) | | `DictObject` | From appPublic.dictObject — available directly (no import) | | `partial` | functools.partial — available directly (no import) | | `FileStorage` | From ahserver.filestorage — available directly (no import) | | `curDateString` / `timestampstr` | From appPublic.timeUtils — date/time string helpers | | `get_config_value(key)` | Get config value | | `exception`, `error`, `debug`, `info`, `warning`, `critical` | Logging functions — all available | | `format_exc` | `traceback.format_exc()` — returns full traceback string (pre-loaded, do NOT `import traceback`) | **Verified via llmage module dspy cleanup (2026-07-01)**: All 31 dspy files had their `import` statements removed and continue to work. The complete list of safely removable imports: `json`, `datetime`, `getID` (appPublic.uniqueID), `debug` (appPublic.log), `curDateString`/`timestampstr` (appPublic.timeUtils), `get_sor_context` (sqlor.dbpools), `time`, `DictObject` (appPublic.dictObject), `partial` (functools), `FileStorage` (ahserver.filestorage), `os`. ## DataViewer CRUD Endpoint Pattern When implementing full CRUD (Create/Update/Delete) for DataViewer widgets, the endpoints must return **Message widget JSON**, not raw data: ```python #!/usr/bin/env python3 # -*- coding: utf-8 -*- """Customer create API for DataViewer editable form""" # No imports needed - json, DBPools, etc. are pre-loaded result = {'widgettype': 'Message', 'options': {'title': 'Error', 'message': 'Invalid request'}} try: name = params_kw.get('customer_name', '') if not name: result['options'] = {'title': 'Error', 'message': 'Name required', 'type': 'error'} else: dbname = get_module_dbname('module_name') async with DBPools().sqlorContext(dbname) as sor: await sor.sqlExe("INSERT INTO table_name (...) VALUES (...)", {...}) result = { 'widgettype': 'Message', 'options': {'title': 'Success', 'message': 'Created successfully', 'type': 'success'} } except Exception as e: result['options'] = {'title': 'Error', 'message': f'Failed: {str(e)}', 'type': 'error'} return json.dumps(result, ensure_ascii=False) ``` **CRITICAL**: The return value MUST be a string (via `json.dumps()`). If the script reaches the end without hitting a `return` statement, ahserver throws `return data type error, `. Every code path must return a string. ## DataViewer Editable Configuration in .ui Configure CRUD operations in the DataViewer's `options.editable` block: ```json { "widgettype": "DataViewer", "options": { "data_url": "/main/module/api/list.dspy", "editable": { "new_data_url": "/main/module/api/create.dspy", "update_data_url": "/main/module/api/update.dspy", "delete_data_url": "/main/module/api/delete.dspy", "form_cheight": 8, "fields": [ {"name": "field_name", "label": "Label", "uitype": "text", "required": true} ] } } } ``` The DataViewer (dataviewer.js) uses these URLs: - `new_data_url` - Form submission URL for adding records - `update_data_url` - Form submission URL for editing records - `delete_data_url` - POST URL for deleting records (sends `{params: row_data}`) ## Common Pitfalls 1. **Using `print()` instead of `return`** — `print()` writes to stdout and is NOT captured by ahserver. The framework expects `return` statements. Using `print()` causes `return data type error, ` because the script returns None. Real-world example: `top_models.dspy` had `print(json.dumps(models))` which returned None to the caller — fixed by changing to `return json.dumps(models, ensure_ascii=False, default=str)`. 2. **Import statements** - violates ahserver security model. All functions listed in the Server-Env table above are pre-loaded — including `getID`, `time`, `DictObject`, `partial`, `FileStorage`, `curDateString`, `timestampstr`. Never import them. If your module's function is needed in a .dspy, export it via `load_{modulename}()` in `init.py`. 3. **Jinja2 `.ui` files cannot execute Python** — `.ui` files are Jinja2 templates that render JSON, they cannot run database queries, async operations, or complex logic. When you need database access, convert to `.dspy` files. Example: `llmusage_ioinfo_display.ui` used `{% set sor = db.sqlorContext() %}` which failed with `NameError: name 'db' is not defined` — fixed by converting to `.dspy` with proper async database access. 4. **sqlPaging() performance pitfall** — `sor.sqlPaging(sql, ns)` wraps the SQL in a subquery `select count(*) from (...)` which is very slow for large tables (3-4 seconds). For better performance, separate count and data queries: ```python # WRONG — sqlPaging is slow for large tables: result = await sor.sqlPaging(sql, ns) # CORRECT — separate count and data queries: count_sql = f"SELECT count(*) as rcnt FROM table {where_clause}" count_recs = await sor.sqlExe(count_sql, ns) total = count_recs[0].rcnt if count_recs else 0 data_sql = f"SELECT ... FROM table {where_clause} ORDER BY {sort} LIMIT {limit} OFFSET {offset}" rows = await sor.sqlExe(data_sql, ns) ``` 5. **Extract reusable database operations to utility functions** — When multiple `.dspy` files need the same database operation (fetching a record, reading from FileStorage), create async functions in the module's `utils.py` and import them: ```python # In module/utils.py: async def get_record_by_id(record_id): env = ServerEnv() async with get_sor_context(env, 'module_name') as sor: sql = "SELECT * FROM table WHERE id = ${id}$" recs = await sor.sqlExe(sql, {'id': record_id}) return dict(recs[0]) if recs else None # In .dspy file: from module.utils import get_record_by_id record = await get_record_by_id(record_id) ``` 6. **FileStorage requires realPath() for file I/O** — FileStorage stores files with webpath references, but actual file operations need the filesystem path: ```python from ahserver.filestorage import FileStorage import aiofiles async def read_storage_file(webpath): fs = FileStorage() real_path = fs.realPath(webpath) # Convert webpath to filesystem path async with aiofiles.open(real_path, 'rb') as f: return await f.read() ``` 3. **Implicit None return** - If any code path doesn't hit a `return` statement, ahserver throws `return data type error, `. **Every branch must end with `return result`.** 4. **CRITICAL: Debug `NoneType` errors at the error location, NOT by adding broad try/except** — When a .dspy endpoint returns `return data type error, `, open the .dspy file FIRST. Do NOT start by adding try/except wrappers in Python functions, modifying database connections, or adjusting SQL. The error trace points directly at the failing .dspy — examine its format (JSON `{"python": {...}}` vs Python script), verify function registration, and check return paths. Broad `except Exception: return []` masks real errors and makes debugging impossible. 5. **JSON-format vs Python-script-format DSPY** — `return data type error, ` is especially common with JSON-format DSPY files (`{"python": {"import": "...", "call": "..."}}`). The JSON-format processor handles `None` returns differently from Python-script format (`import json; data = await func(request); return json.dumps(data)`). If one DSPY in a module uses JSON format while all others use Python script format, it's likely a format inconsistency bug. Always check file format when debugging NoneType errors. This is the single most common dspy error — the dspy sets `result` in branches but forgets the final `return result` at module level. Even a trivial dspy like `result = {"text": "hello"}` will return None without an explicit `return result`. **Pattern for multi-branch dspy**: put `return result` at the very end, OUTSIDE all if/elif blocks: ```python if not user: result = {...} elif code: result = {...} else: result = {...} return result # ← REQUIRED, outside all branches ``` 4. **Wrong data format** - code components need `{value, text}` arrays 5. **Missing error handling** - causes 500 errors instead of graceful degradation 6. **Returning wrapper objects unnecessarily** - most components expect direct data 7. **SQL column mismatch with DDL** - SELECT columns in .dspy files MUST exactly match actual database schema. Always verify with `DESCRIBE table_name` before writing queries. DDL files may differ from deployed schema. 8. **CGI-style .dspy files** — Never use `os.environ`, `sys.stdin`, `os.read(0, ...)`, `print()`, or `asyncio.new_event_loop()`. Use `params_kw`, `sqlorContext`, `return`, and let ahserver handle the async context. **ahserver automatically parses ALL request data (query string + POST body, including JSON `application/json`) into `params_kw`** — no manual reading of stdin or `os.read(0, content_length)` needed. JSON POST bodies are preserved as nested dict/list structures: `params_kw.get('user', {})` returns the nested user object. ❌ Never write `content_length = int(os.environ.get('CONTENT_LENGTH', 0)); raw_data = os.read(0, content_length); post_data = json.loads(raw_data)` in a .dspy file. 9. **DataViewer CRUD endpoints returning raw JSON** - Create/update/delete endpoints called by DataViewer editable forms must return a `Message` widget JSON structure, not raw data dictionaries. 10. **Dict merge syntax `{**a, **b}` fails** - Use `dict(list(a.items()) + list(b.items()))` instead for merging parameter dictionaries in .dspy files. 11. **sqlExe return type depends on parameters** - `sor.sqlExe(sql, ns)` returns different types: - **WITHOUT `page`/`rows` in ns**: returns a **list** of row objects — do NOT treat as dict (`ret['key']` will fail with TypeError) - **WITH `page`/`rows` in ns**: returns a **dict** `{'total': N, 'rows': [...]}` — do NOT iterate directly as list Always check type or build result manually: ```python rows = await sor.sqlExe(sql, ns) # no page/rows result = {'total': len(rows), 'rows': rows, 'stats': stats} return json.dumps(result, ensure_ascii=False, default=str) ``` 12. **Row objects need safe conversion** - sqlExe returns row objects that may or may not have `.keys()` method. Use `dict(r) if hasattr(r, 'keys') else r` for safe conversion. 13. **Sort column must exist in table** - sqlExe uses the `sort` parameter for ORDER BY. If the specified column doesn't exist in the table, query fails. Default 'id' may not always be available. 14. **API file location matters** - List API `.dspy` files must be in `wwwroot/api/` subdirectory (e.g., `wwwroot/api/customers_list.dspy`), while UI `.ui` files go directly in `wwwroot/`. 15. **Session expiration during testing** - Cookie sessions expire after `session_max_time` (default 3600s). Re-login via `/main/login.dspy?username=xxx&password=xxx` before testing if getting 401 errors. 22. **Connection pool dirty reads (multiserver)**: When multiple service instances share the same MySQL, a connection pool bug can cause `get_*.dspy` to read records that were just deleted by another instance. Root cause: `aiomysql.connect()` defaults to `autocommit=False`, and `sqlorContext` only calls `commit()` for writes (not reads). When a connection is reused from the pool, its REPEATABLE READ snapshot from a previous SELECT persists. Fix: in `mysqlor.enter()`, call `await self.conn.commit()` to end any lingering transaction before reuse. See sqlor repo commit `fab420c`. Symptom: high-frequency "get reads deleted record" reports in multi-instance deployments. 24. **CRITICAL: Filter NaN/null/empty before MySQL INSERT/UPDATE** — When receiving numeric parameters from bricks `UiFloat` widgets via `urlwidget` + `datawidget: "self"`, empty or invalid inputs may send `NaN`, `null`, or empty strings. MySQL cannot handle `nan` floats: `OperationalError: nan can not be used with MySQL`. Always sanitize: ```python discount_val = params_kw.get('discount') if discount_val is not None: s = str(discount_val).strip().lower() if s in ('', 'nan', 'none', 'null'): discount_val = None ``` Apply this to ALL numeric parameters from user input before `float()` conversion or sor.C/U. Also applies to `old_discount` comparison values sent as static params. 23. **CRITICAL: SQL parameter syntax** - Use `${param}$` in SQL strings, NOT `%(param)s`. The `${param}$` placeholder is replaced by sqlor with proper escaping. Using `%(param)s` causes "format requires a mapping" errors. Example: `await sor.sqlExe("INSERT INTO t (col) VALUES (${col}$)", {'col': value})`. 17. **Optional DATE fields** - MySQL DATE columns reject empty strings `''`. Convert empty form values to `None`: `sign_date = params_kw.get('sign_date', '').strip() or None`. 18. **Safe row attribute access** - SQLor row objects may lack `.keys()` or dict access. Use `getattr(row, 'field', '') or ''` instead of `row.field` or `row['field']` to avoid AttributeError on missing/None fields. 19. **$or batch queries** - For looking up multiple records by ID, build `$or` conditions: `{'$or': [{'id': uid} for uid in ids]}`. Validate the ID list is non-empty before querying. 20. **CRITICAL: ServerEnv() forbidden in .dspy AND in Python helper functions** — The ahserver framework injects all necessary functions directly into the .dspy execution context via globals. **Never write `env = ServerEnv()` in a .dspy file.** Correct usage: `dbname = get_module_dbname('dapi')`, `db = DBPools()`, `create_apikey_func = create_user_apikey`. Using `getattr(env, 'func_name', None)` or `config = getConfig(); db.databases = config.databases` is also wrong — these are all available as bare names. **For Python functions in `init.py` called from dspy**: Use `env = request._run_ns`, NOT `env = ServerEnv()`. A bare `ServerEnv()` has no request binding — `get_user()`, `get_userorgid()`, `get_userid()` etc. will all be `None`. The correct pattern: ```python # ✅ CORRECT — request._run_ns has full request context async def my_handler(request, params_kw): env = request._run_ns user_id = await env.get_user() # returns userid string org_id = await env.get_userorgid() # returns orgid string # ❌ WRONG — bare ServerEnv() has no session/request binding async def my_handler(request, params_kw): env = ServerEnv() user_id = await env.get_user() # None! 'NoneType' is not callable ``` **Symptom**: dspy returns 500, log shows `'NoneType' object is not callable` at calls like `env.get_userorgid()` or `env.get_user()`. 25. **Bare function calls from `load_X()` registrations can be None in dspy context** — Functions registered via `env.func_name = func` in `load_discount()` (etc.) are placed on the `ServerEnv` singleton, which gets merged into the dspy execution namespace via `run_ns.update(ServerEnv())`. In practice, this merge can fail silently — the bare function name resolves to `None` in the dspy, producing `'NoneType' object is not callable`. When a bare function call returns this error, **use `request._run_ns.func()` instead of bare function calls:** ```python # ❌ Bare function call — may resolve to None in dspy context: file_type = classify_file(file_name) # NameError or NoneType # ❌ Explicit import — PROHIBITED in .dspy (user-enforced rule): from rag.pipeline import process_upload # BLOCKED # ❌ `request._run_ns.func()` — PROVEN UNRELIABLE in production (#6) # ServerEnv registration in init_rag_module() does NOT propagate to DSPY exec context. # Despite env.func = func being set correctly, request._run_ns.func is always None. # # ✅ THE ONLY RELIABLE PATTERN — inline all logic directly in the DSPY: # Use only ahserver pre-loaded globals: json, uuid, DBPools, get_sor_context, # request.read(), params_kw. For installed packages (PyPDF2, docx, pptx, openpyxl), # import inline at point of use — these are venv-installed, not custom modules. env = request._run_ns result = await env.process_upload(env, file_data, kb_id, folder_id, file_name) return result ``` **Registration in init.py** (module's `init_rag_module` or `load_rag`): ```python def init_rag_module(): env = ServerEnv() from .pipeline import process_upload env.process_upload = process_upload rf = RegisterFunction() ... ``` **DSPY becomes a zero-import thin wrapper** (12 lines max): ```python ns = params_kw.copy() kb_id = ns.get('kb_id', '') folder_id = ns.get('folder', '') file_name = ns.get('file_name', 'upload.bin') if not kb_id: return json.dumps({"status": "error", "error": "kb_id required"}, ensure_ascii=False) file_data = await request.read() if not file_data: return json.dumps({"status": "error", "error": "no file data"}, ensure_ascii=False) env = request._run_ns result = await env.process_upload(env, file_data, kb_id, folder_id, file_name) return result ``` This pattern was verified on ragserver (yumoqing/rag.git) — bare function calls and explicit imports both fail; only `request._run_ns.func()` works reliably. Keep all business logic in Python modules (pipeline.py, utils.py); DSPY files are pure wire-up. ```python # ❌ May resolve to None in dspy context: ret = await bind_customer(request, bind_params) # NoneType not callable # ✅ Reliable — explicit import bypasses namespace merge issues: from discount.init import bind_customer, set_promote_discount ret = await bind_customer(request, bind_params) # works ``` **Diagnosis**: create a minimal test .dspy: `result = {'text': str(type(bind_customer))}` — if output shows ``, the function isn't being found in the namespace. **VERIFIED DECISION (ragserver, 2026-07-29)**: ALL approaches were tested exhaustively: 1. `env.func = func` in `init_rag_module()` → `request._run_ns.func` always None in DSPY exec context 2. `from rag.pipeline import func` → blocked by user (imports not allowed in DSPY) 3. **Inline all logic directly in the DSPY** → the ONLY approach that works Use only ahserver pre-loaded globals (`json`, `uuid`, `DBPools`, `get_sor_context`, `request.read()`, `params_kw`, `request._run_ns.get_userorgid()`). For installed packages (`PyPDF2`, `docx`, `pptx`, `openpyxl`, `aiohttp`, `base64`), import inline at point of use — these are venv-installed packages, NOT custom module imports. The DSPY file becomes a self-contained script with zero custom imports. For building complete upload pipelines with text extraction + DB, put ALL logic in the DSPY — do NOT attempt to split across pipeline.py or module init.py. 26. **CRITICAL: f-string braces inside dict returns cause exec() parse error** — `exec()` interprets f-string `{e}`'s closing `}` as closing the outer dict, producing `SyntaxError: '{' was never closed`. ```python # ❌ exec() misreads the last } — thinks it closes the outer dict return {"timeout": 5, "message": f"处理失败: {e}"} # ✅ Use string concatenation instead return {"timeout": 5, "message": "处理失败: " + str(e)} ``` This also affects `exception(f'{var=}')` — the `=` inside `{var=}` is fine but the closing `}` before `)` triggers the same issue. Use `'prefix: ' + str(var)` for debug/exception calls too. — `sor.sqlExe(sql, ns)` without page/rows returns a list of **row objects** (like SimpleNamespace), not dictionaries. These objects support attribute access (`r.id`, `r.name`) but NOT dict access (`r['id']`, `r['name']`). Using dict access causes `TypeError: 'SimpleNamespace' object is not subscriptable`, which can be silently swallowed by `try/except` blocks, resulting in empty dropdowns or undefined values in the UI. **❌ Wrong (causes silent failure):** ```python apps = await sor.sqlExe("select id, name from upapp", {}) result = [{'value': r['id'], 'text': r['name']} for r in apps] # TypeError silently caught ``` **✅ Correct:** ```python apps = await sor.sqlExe("select id, name from upapp", {}) result = [{'value': str(r.id), 'text': r.name} for r in apps] # Attribute access ``` **Safe pattern with getattr:** ```python apps = await sor.sqlExe("select id, name from upapp", {}) result = [{'value': str(getattr(r, 'id', '')), 'text': getattr(r, 'name', '')} for r in apps] ``` **Why this matters:** When building dropdown data endpoints (like `get_upapps.dspy`), using dict access causes the endpoint to return an empty array `[]`, which makes dropdown fields show "undefined" in the UI. The error is invisible because the try/except catches it silently. ## Module Deployment Workflow **CRITICAL**: Never edit code directly on test/production servers. All changes must follow this flow: 1. Edit in local repo (`~/repos//`) 2. `git add` + `git commit` + `git push` 3. On test server: `git pull` in the module's directory 4. If server has no SSH key for git, scp changed files individually **Module directory structure** (Sage): ``` /d/apitest/sage/ pkgs/ module_name/ ← git repo (for code) wwwroot/ ← symlinked from ../../wwwroot/module_name module_name/ ← Python package (copied to site-packages) wwwroot/ module_name -> ../pkgs/module_name/wwwroot ← symlink py3/lib/python3.10/site-packages/ module_name/ ← Python package (copied from pkgs during deploy) ``` Modules live under Sage's `pkgs/` directory, NOT under pipeline-app's `pkgs/`. Each module's `wwwroot/` is symlinked from Sage's main `wwwroot/`. Python code is copied to `site-packages/` for the Sage venv to find. When you cannot run direct database queries, create a temporary debug `.dspy` file to inspect table schemas: ```python # Debug: show table columns — no imports needed result = {'keys': [], 'rows': []} try: dbname = get_module_dbname('module_name') async with DBPools().sqlorContext(dbname) as sor: ns = {'page': 1, 'rows': 50, 'sort': 'COLUMN_NAME'} sql = "SELECT COLUMN_NAME, COLUMN_TYPE FROM information_schema.COLUMNS WHERE TABLE_SCHEMA='dbname' AND TABLE_NAME='table_name'" rows = await sor.sqlExe(sql, ns) if isinstance(rows, dict): rows = rows.get('rows', []) if rows: result['keys'] = list(dict(rows[0]).keys()) result['rows'] = [list(dict(r).values()) for r in rows] result['success'] = True except Exception as e: result['error'] = str(e) return json.dumps(result, ensure_ascii=False, default=str) ``` Place in `wwwroot/api/debug_tables.dspy`, test via curl, then delete after getting schema info. ## Best Practices Summary - ✅ Use `return` for all data responses - ✅ Never use `import` statements - ✅ Handle all exceptions gracefully - ✅ Return component-appropriate data formats - ✅ Validate all input parameters - ✅ Keep .dspy files focused and minimal - ✅ Only create .dspy files when standard CRUD endpoints are insufficient - ✅ Follow consistent naming patterns (`/list/`, `/get/`, `/test/`, etc.) - ✅ Use `getattr(row, 'field', '') or ''` for safe SQLor row attribute access - ✅ Use bare `get_module_dbname('module')` for cross-module DB access (no `ServerEnv()` wrapper needed) - ✅ Use `$or` conditions in sor.R for batch ID lookups ## Complex Logic: Move to Python, DSPY as Thin Wrapper When a .dspy needs to call module-internal classes or functions not registered on ServerEnv (e.g., `EmailClient`, `PROVIDERS`, provider methods), the DSPY will hit `NameError`. **Never add imports to the DSPY.** Instead: 1. Add the logic as a method on a provider class (e.g., `TransferGateway.check_transfer()`) 2. Register the provider on ServerEnv: `env.PROVIDERS = PROVIDERS` 3. The DSPY becomes a thin wrapper: ```python provider = env.PROVIDERS.get('transfer') title, msg = await provider.check_transfer(tcode, env) return {"widgettype": "Message", "options": {"title": title, "message": msg}} ``` **This also avoids f-string brace issues** (pitfall 26) — the Python method can use f-strings freely; only the DSPY wrapper uses concatenation. ## add_startup Blocks Server — Use Manually Triggered Actions `add_startup(coro)` awaits the coroutine during server startup. If the coroutine is an infinite `while True` loop, **it blocks the server indefinitely**. Never use `add_startup` with an infinite loop or long-running polling. Instead, trigger actions manually (e.g., a button calling a DSPY endpoint) or use `asyncio.create_task()` inside the startup callback to spawn non-blocking background tasks. ## How DSPY Execution Works (ahserver wraps in async function) **CRITICAL**: The ahserver framework wraps your .dspy code in an async function and awaits it: ```python # ahserver baseProcessor.py line ~234-243: txt = "async def myfunc(request,**ns):\n" + '\n'.join(lines) exec(txt, lenv, lenv) func = lenv['myfunc'] return await func(request, **lenv) ``` This means: - `async with`, `await`, and `async for` DO work inside .dspy files - You MUST use explicit `return` — the function's return value is what gets passed to the caller - A bare expression (like `result` on the last line) inside an `async with` block does NOT reach the outer scope — it's local to the async function **❌ WRONG — bare expression, function returns None:** ```python async with db.sqlorContext(dbname) as sor: data = await sor.R('table', {}) result = [dict(r) for r in data] result # ← local to function, not returned ``` **✅ CORRECT — explicit return:** ```python async with db.sqlorContext(dbname) as sor: data = await sor.R('table', {}) return [dict(r) for r in data] return [] ``` This pattern is used extensively in the RBAC permission CRUD dspy files (e.g., `get_permission.dspy`) and our `get_tree_data.dspy` / `new_tree_item.dspy`, all of which work correctly. ## Async/Await — Fully Supported in DSPY **VERIFIED (2026-07-29, ragserver)**: `async with`, `await`, and `async for` ALL work inside .dspy files. The ahserver framework wraps your code in `async def myfunc(request, **ns):` and awaits it. The earlier prohibition was incorrect — extensive testing on the ragserver module confirmed all async patterns work: ```python # ✅ ALL of these work in DSPY: async with get_sor_context(env, 'rag') as sor: recs = await sor.sqlExe("SELECT ...", {}) file_data = await request.read() async with aiohttp.ClientSession() as s: r = await s.post('https://...', json={...}) ``` **Symptom of real async issues**: 500 with `'NoneType' object is not callable` — this is almost always a ServerEnv registration failure (see Pitfall 25), NOT an async/sync problem. The function is None, not uncallable because of async context. **PITFALL: `params_kw` unavailable in some DSPY contexts** — when a `.dspy` file is accessed as a standalone page endpoint (like `/discount/promote.dspy`), `params_kw` may not be in scope. Use `request._run_ns.params_kw` instead: ```python # ✅ Safe — works in all DSPY contexts code = request._run_ns.params_kw.get('code', '') # ❌ May fail — params_kw not always available code = params_kw.get('code', '') ``` **PITFALL: `binds` with `script` actiontype causes 500 in DSPY files** — when a DSPY returns widget JSON containing a `binds` array with `actiontype: "script"`, the server-side JSON parser may attempt to evaluate the script string as Python, causing 500 errors. Avoid including `binds` in DSPY widget output; keep them in static `.ui` templates instead. async with DBPools().sqlorContext(dbname) as sor: recs = await sor.R('discount_promo_code', {'id': promo_id}) ... ``` **Pattern**: export the async function via `load_discount()` (`env.generate_promo_qr = generate_promo_qr`), then call it from the DSPY with `await func_name(request, params_kw)`. The DSPY stays a thin 2-line wrapper. **Symptom**: access to `.dspy` returns 500, server log shows `return data type error, ` and `'NoneType' object is not callable` from `auth_api.py`. **Note**: This prohibition does NOT apply to CRUD wrapper DSPY files (see "CRUD Wrapper Pattern" below) — those wrappers use `await` legitimately because they delegate to pre-registered async functions. ## CRUD Wrapper Pattern (Legitimate Exception) When a module's `init.py` registers CRUD functions via `load_{module}()` (e.g., `env.create_tablename = create_tablename`), the `wwwroot/api/*.dspy` files are **thin wrappers** that delegate to those functions. These wrappers use `ServerEnv()` and `print()` — this is a legitimate exception to the "no ServerEnv in dspy" rule. **CRITICAL**: `json` is pre-loaded in ALL dspy contexts (including wrappers). Do NOT `import json` — it is redundant and will cause pre-commit audit failures. The only import needed is `from ahserver.serverenv import ServerEnv`: ```python from ahserver.serverenv import ServerEnv env = ServerEnv() create_func = getattr(env, 'create_tablename', None) if create_func is None: print(json.dumps({"status": "error", "message": "create_tablename function not found"})) else: result = await create_func(request, params_kw) print(result) ``` **When this pattern applies**: Only for `wwwroot/api/{table}_create.dspy`, `{table}_update.dspy`, `{table}_delete.dspy` files that delegate to init.py-registered CRUD functions. **When NOT to use**: Business logic .dspy files that do actual work (queries, calculations, cross-module operations) must follow the standard pattern (no imports, no ServerEnv, use return). ## Linked References - `references/user-sync-pattern.md` — Cross-module user sync API pattern - `references/dirty-apikey-record-pattern.md` — Orphan downapikey records - `references/accounting-table-architecture.md` — Accounting table schema - `references/sage-deploy-test-server.md` — Sage module deployment workflow - `references/pipeline-app-setup.md` — Pipeline-app config, debugging, and KTV setup - `references/sage-crontab-etl-pattern.md` — Cron DSPY endpoints + build.sh crontab + j2_ stat cards - `references/cross-table-column-pitfalls.md` — Column name mismatches across Sage tables (userorgid vs orgid) + catelogid length + GROUP BY ambiguity - sqlor-database-module skill `references/dapi-table-architecture.md` — Full dapi module table structure