nomyo-router

Author	SHA1	Message	Date
alpha-nerd-nomyo	fbdc73eebb	fix: improvements, fixes and opt-in cache doc: semantic-cache.md added with detailed write-up	2026-03-10 15:19:37 +01:00
alpha-nerd-nomyo	dd4b12da6a	feat: adding a semantic cache layer	2026-03-08 09:12:09 +01:00
alpha-nerd-nomyo	b951cc82e3	bump version	2026-03-05 11:09:20 +01:00
alpha-nerd-nomyo	8037706f0b	fix(db.py): remove full table scans with proper where clauses for dashboard statistics and calc in db rather than python	2026-03-03 17:20:33 +01:00
alpha-nerd-nomyo	45315790d1	fix(router.py): - added global for orphaned token_worker_task and flust_task - fixed a regex to effectively _mask_secrets - fixed several Type and KeyErrors - fixed model deduplication for llama_server_endpoints	2026-03-03 16:34:16 +01:00
alpha-nerd-nomyo	e96e890511	refactor: make choose_endpoint use cache incrementer for atomic updates	2026-03-03 14:57:37 +01:00
alpha-nerd-nomyo	10c83c3e1e	fix(router): treat missing status as loaded for llama model check Add check for `status is None` in `_is_llama_model_loaded`. Models without a status field (e.g., single-model servers) are assumed to be always loaded rather than failing the check. Also updated docstring to clarify this behavior.	2026-03-02 08:54:46 +01:00
alpha-nerd-nomyo	cac0580eec	feat: adding /v1/rerank endpoint with cohere,jina,llama.cpp compatibility	2026-02-28 09:31:25 +01:00
alpha-nerd-nomyo	ad4a1d07b2	fix(/v1/embeddings): returning the async_gen forced FastAPI serialization which caused Pydantic Errors. Also sanizted nan/inf values to floats (0.0). Use try - finally to properly decrement usage counters in case of error.	2026-02-27 16:39:27 +01:00
alpha-nerd-nomyo	d2ea65f74a	fix(router): use normalized model keys for endpoint selection Refactor endpoint selection logic to consistently use tracking model keys (normalized via `get_tracking_model`) instead of raw model names, ensuring usage counts are accurately compared with how increment/decrement operations store them. This fixes inconsistent load balancing and model affinity behavior caused by mismatches between raw and tracked model identifiers.	2026-02-19 17:32:54 +01:00
alpha-nerd-nomyo	07751ddd3b	fix: endpoint selection logic again	2026-02-19 10:11:53 +01:00
alpha-nerd-nomyo	7cba67cce0	feat(router): normalize model names for usage tracking across endpoints (continued) Introduce `get_tracking_model()` to standardize model names for consistent usage tracking in Prometheus metrics. This ensures llama-server models are stripped of HF prefixes and quantization suffixes, Ollama models append `:latest` when versionless, and external OpenAI models remain unchanged—aligning all tracking keys with the PS table.	2026-02-18 11:45:37 +01:00
alpha-nerd-nomyo	b2980a7d24	fix(router): handle invalid version responses with 503 error Filter out non-string version responses (e.g., empty lists from failed requests) and return a 503 Service Unavailable error if no valid versions are received from any endpoint.	2026-02-17 15:56:09 +01:00
alpha-nerd-nomyo	836c5f41ea	fix(router): normalize model names for usage tracking across endpoints	2026-02-17 11:35:53 +01:00
alpha-nerd-nomyo	372fe9fb72	feat(router): parallelize llama-server props fetch and add reasoning/tool call support - Fetch `/props` endpoints in parallel to get context length and auto-unload sleeping models - Add support for reasoning content and tool calls in streaming openai chat/completions responses	2026-02-15 17:05:35 +01:00
alpha-nerd-nomyo	4d40048fd2	fix: loaded_models_cache timing restored	2026-02-15 12:15:36 +01:00
alpha-nerd-nomyo	0bad604b02	feat: deduplicate background refresh tasks and extend cache TTL Adds lock-protected dictionaries to track running background refresh tasks, preventing duplicate executions per endpoint. Increases cache freshness thresholds from 30s to 300s to reduce blocking behavior. fix: /v1 endpoints use correct media_types and usage information with proper logging	2026-02-14 14:51:44 +01:00
alpha-nerd-nomyo	c9ff384bb2	fix(router): /v1/models endpoint Shows now all available models	2026-02-13 16:27:06 +01:00
alpha-nerd-nomyo	4d80dc5e7c	feat: adding logprobs to /v1/chat/completion	2026-02-13 14:43:10 +01:00
alpha-nerd-nomyo	eda48562da	feat(router): add logprob support in /api/chat Add logprob support to the OpenAI-to-Ollama proxy by converting OpenAI logprob formats to Ollama types. Also update the ollama dependency.	2026-02-13 13:29:45 +01:00
alpha-nerd-nomyo	1b355d8435	Merge branch 'main' into dev-v0.6	2026-02-13 10:33:36 +01:00
alpha-nerd-nomyo	08b77428b8	refactor(router): bump cache TTLs and skip error cache for health checks - Increased error and loaded model cache freshness thresholds from 10s to 30s. - Added `skip_error_cache` parameter to `endpoint_details` to prevent cached failures from blocking health checks. - Implemented automatic error recording in `_available_error_cache` on API request failures.	2026-02-13 10:11:41 +01:00
alpha-nerd-nomyo	b649dcd8d6	proposal: use global truststore ctx for all connections	2026-02-12 16:15:39 +01:00
Jan-Timo	dd30ab9422	fix SSL: CERTIFICATE_VERIFY_FAILED	2026-02-11 13:47:11 +01:00
alpha-nerd-nomyo	9875eb977a	feat: Add tool call normalization and streaming delta accumulation Adds support for correctly handling tool calls in chat requests. Normalizes tool call data (ensuring IDs, types, and JSON arguments) in non-streaming mode and accumulates OpenAI-style deltas during streaming to build the final Ollama response.	2026-02-10 20:21:46 +01:00
alpha-nerd-nomyo	4892998abc	feat(router): Add llama-server endpoints support and model parsing Add `llama_server_endpoints` configuration field to support llama_server OpenAI-compatible endpoints for status checks. Implement helper functions to parse model names and quantization levels from llama-server responses (best effort). Update `is_ext_openai_endpoint` to properly distinguish these endpoints from external OpenAI services. Update sample configuration documentation.	2026-02-10 16:46:51 +01:00
alpha-nerd-nomyo	1f81e69ce1	refactor(router.py): correctly implement OpenAI tool_calls to Ollama format conversion	2026-02-09 11:04:14 +01:00
alpha-nerd-nomyo	7deb088c6a	refactor(cache): split error cache and add stale-while-revalidate Refactor error tracking to use separate caches for 'available' and 'loaded' models, preventing cross-contamination of transient errors. Implement background refresh for available models to prevent blocking requests, and use stale-while-revalidate (300-600s) to serve stale data immediately when the cache is between 300s and 600s old.	2026-02-08 16:46:40 +01:00
alpha-nerd-nomyo	92cea1dead	feat: update reasoning handling Updated reasoning content handling in router.py to check for both "reasoning_content" and "reasoning" attributes.	2026-02-08 11:29:47 +01:00
alpha-nerd-nomyo	bd0d210b2a	feat: enforce api key authentication and update table header - Added proper API key validation in router.py with 401 response when key is missing - Implemented CORS headers for authentication requests - Updated table header from "Until" to "Unload" in static/index.html - Improved security by preventing API key leakage in access logs	2026-02-01 10:05:46 +01:00
alpha-nerd-nomyo	b718d575b7	Merge pull request #22 from nomyo-ai/dev-v0.5.x Dev v0.5.x	2026-01-30 18:18:42 +01:00
alpha-nerd-nomyo	d80b29e4f2	feat: enhance code quality and documentation - Renamed Feedback class to follow PascalCase convention - Fixed candidate enumeration start index from 0 to 1 - Simplified candidate content access by removing .message.content - Updated CONFIG_PATH environment variable name to CONFIG_PATH_ARG - Bumped version from 0.5 to 0.6 - Removed unnecessary return statement and trailing newline	2026-01-29 19:59:08 +01:00
alpha-nerd-nomyo	4ca1a5667e	feat(router): implement in-flight request tracking to prevent cache stampede in high concurrency scenarios Added in-flight request tracking mechanism to prevent cache stampede when multiple concurrent requests arrive for the same endpoint. This introduces new dictionaries to track ongoing requests and a lock to coordinate access. The available_models method was refactored to use an internal helper function and includes request coalescing logic to ensure only one HTTP request is made per endpoint when cache entries expire. The loaded_models method was also updated to use the new caching and coalescing pattern.	2026-01-29 18:00:33 +01:00
alpha-nerd-nomyo	a1276e3de8	fix: correct indentation for publish_snapshot calls in usage functions This fix ensures that the snapshot publishing happens within the usage lock context, maintaining proper synchronization of usage counts.	2026-01-29 10:32:59 +01:00
YetheSamartaka	d3aa87ca15	Added endpoint differentiation for models ps board Added endpoint differentiation for models PS board to see where which model is loaded and for how long to ease the viewing of multiple same models deployed for load balancing	2026-01-27 13:29:54 +01:00
alpha-nerd-nomyo	ee1c460477	Empty key strings could bypass authentication in _extract_router_api_key() when malformed Authorization headers were sent - Added validation to check that the extracted key is not empty before returning it - Added CORS headers to enforce_router_api_key() for proper cross-origin request handling and CORS-related error prevention	2026-01-26 18:11:28 +01:00
alpha-nerd-nomyo	d4b2558116	refactor: improve snapshot safety and usage tracking Create atomic snapshots by deep copying usage data structures to prevent race conditions. Protect concurrent reads of usage counts with explicit locking in endpoint selection. Replace README screenshot with a video link.	2026-01-26 17:18:57 +01:00
alpha-nerd-nomyo	3e3f0dd383	fix: endpoint selection logic	2026-01-19 14:21:08 +01:00
alpha-nerd-nomyo	5ad5bfe66e	feat: endpoint selection more consistent and understandable	2026-01-18 09:31:53 +01:00
alpha-nerd-nomyo	067cdf641a	feat: add timestamp index and improve cache concurrency - Added index on token_time_series timestamp for faster queries - Introduced cache locks to prevent race conditions	2026-01-16 16:47:24 +01:00
YetheSamartaka	eca4a92a33	add: Optional router-level API key that gates router/API/web UI access Optional router-level API key that gates router/API/web UI access (leave empty to disable) ## Supplying the router API key If you set `nomyo-router-api-key` in `config.yaml` (or `NOMYO_ROUTER_API_KEY` env), every request to NOMYO Router must include the key: - HTTP header (recommended): `Authorization: Bearer <router_key>` - Query param (fallback): `?api_key=<router_key>` Examples: ```bash curl -H "Authorization: Bearer $NOMYO_ROUTER_API_KEY" http://localhost:12434/api/tags curl "http://localhost:12434/api/tags?api_key=$NOMYO_ROUTER_API_KEY" ```	2026-01-14 09:28:02 +01:00
alpha-nerd-nomyo	20a016269d	feat: added buffer_lock to prevent race condition in high concurrency scenarios added documentation	2026-01-05 17:16:31 +01:00
alpha-nerd-nomyo	19a13cc613	fix(enhance.py): correct typo in function name from 'moe_select_candiadate' to 'moe_select_candidate' feat(router.py): add helper function _make_chat_request for handling enhancing chat requests to endpoints	2025-12-15 10:35:56 +01:00
alpha-nerd-nomyo	5eb5490d16	feat: improve model version handling in endpoint selection Add logic to only append ":latest" suffix to models without existing version suffixes, preventing duplicate version tags and ensuring correct endpoint selection for models following Ollama naming conventions.	2025-12-14 17:58:45 +01:00
alpha-nerd-nomyo	3ccaf78e5d	fix: simplify model version handling in proxy functions Simplify the logic for handling model versions in `openai_chat_completions_proxy` and `openai_completions_proxy` by removing redundant conditions and initializing `local_model` earlier. This makes the code more readable and maintains the same functionality.	2025-12-13 12:34:24 +01:00
alpha-nerd-nomyo	34d6abd28b	refactor: optimize token aggregation query and enhance chat proxy - Refactored token aggregation query in db.py to use a single SQL query with SUM() instead of iterating through rows, improving performance - Combined import statements in db.py and router.py to reduce lines of code - Enhanced chat proxy in router.py to handle "moe-" prefixed models with multiple query execution and critique generation - Added last_user_content() helper function to extract user content from messages - Improved code readability and maintainability through these structural changes	2025-12-13 11:58:49 +01:00
alpha-nerd-nomyo	59a8ef3abb	refactor: use a persistent WAL-enabled connection with async locks - Introduce a lazily initialized, shared aiosqlite connection stored in self._db and two asyncio locks (_db_lock, _operation_lock) for safe concurrent access - Ensure the database directory exists before connecting and enable WAL journaling and foreign keys on first connect - Add close method to gracefully close the persistent connection - Guard initialization and write operations with _operation_lock to ensure single-threaded schema setup - Switch to ON CONFLICT UPSERT for token_counts updates and initialize token_time_series table - Add typing for _db (Optional[aiosqlite.Connection]) and adjust imports accordingly addition: Frontend button with total stats aggregation task and feedback span element to keep user informed and a small database footprint	2025-12-02 12:18:23 +01:00
alpha-nerd-nomyo	0ffb321154	fixing total stats model, button, labels and code clean up	2025-11-28 14:59:29 +01:00
alpha-nerd-nomyo	1c3f9a9dc4	fix model naming to allow correct decrement usage counter in /v1 endpoints	2025-11-24 09:33:54 +01:00
alpha-nerd-nomyo	7b50a5a299	adding usage metrics to /v1 endpoints if stream == True	2025-11-21 09:56:42 +01:00

1 2 3

118 commits