11 KiB
11 KiB
Nexus One AI — Current Capability Catalog
Last verified: 21 July 2026
Reference environment: 172.16.10.144
Portal: https://172.16.10.144:9445
Open WebUI: https://172.16.10.144:9446
This document is the presentation and product-description source of truth for what each visible Nexus One AI tool does. “Certified” means its core action was completed against the deployed backend with a real persisted result. It does not mean every scale, hardware, security, or production-hardening scenario has been certified.
Product in one sentence
Nexus One AI is a customer-controlled enterprise AI platform that combines local model inference, governed knowledge and automation workflows, evaluation, application development, Kubernetes operations, monitoring, and recovery behind one identity and portal.
User-facing AI tools
| Tool | What it does | Current verified result |
|---|---|---|
| AI Workspace | Provides persistent multi-conversation chat with installed local models, system prompts, temperature controls, and image input where supported. | Certified with real local inference; failed model calls are shown as failures and do not fall back to simulated answers. |
| Nexus Notebook / Knowledge Base | Uploads and indexes documents, searches ChromaDB, and answers questions using retrieved source passages. | Certified with document ingestion, vector retrieval, grounded answers, and citations. |
| Meeting Assistant | Converts a transcript into a summary, topics, decisions, sentiment, action items, and formatted minutes. | Transcript-to-minutes flow certified with real model output. Audio transcription remains dependent on supported audio and speech-model availability. |
| Prompt Library | Stores reusable prompts by business category, tags, and document requirements. | Certified for authenticated listing and administrative create, update, and delete operations. |
| Prompt Studio | Builds prompts with personas, system instructions, parameters, optional document context, history, and hand-off to comparison or evaluation. | Certified with a seeded executive demo prompt and real model execution. |
| Response Comparison | Sends the same system and user prompt to two to four installed models and displays response, latency, and token usage side by side. | Certified with genuine phi3:mini and llama3.1:8b responses. Browser history is presentation convenience, not a server-side evaluation record. |
| Agent Studio | Defines governed multi-step agents that call models and platform tools. | Certified with a completed real agent pipeline and persisted execution result. |
| Workflow Designer | Executes ordered AI, retrieval, transformation, and integration steps. | Certified with a real backend AI step; backend failure is surfaced and never replaced by simulated completion. |
| Scheduled Jobs | Runs approved agent or workflow actions on a schedule or through a manual “Run now” action and stores the result. | Certified with a real manually triggered scheduled job. |
| AI Evaluation Suite | Runs repeatable test cases through selected models and records model and judge scores for quality, relevance, and safety. | Certified with actual generation and persisted evaluation scores. |
| RAG Quality | Measures retrieval relevance and grounded-answer quality against a selected knowledge collection. | Core page and collection integration are available; quality depends on the indexed source material and evaluation set. |
| Feedback Review | Captures user ratings and comments, supports administrative review, and promotes useful examples into evaluation workflows. | Certified through feedback review and evaluation integration. |
| Chat Rooms | Provides shared rooms for human discussion with @AI participation using local models. |
Certified with authenticated room messaging and a real @AI response. |
| Dataset Creation | Converts PDF, DOCX, TXT, Markdown, or CSV sources into grounded supervised examples, supports review, and publishes approved rows as JSONL. | Certified end to end: five grounded examples generated with llama3.1:8b, reviewed, approved, and published. Jobs now fail honestly if the requested valid sample count is not reached. |
| Model Training | Fine-tunes a supported Hugging Face causal language model with a validated dataset and saves a PEFT LoRA adapter, tokenizer, configuration, logs, loss, and output path. | Certified on CUDA with HuggingFaceTB/SmolLM2-135M-Instruct, five examples, one epoch, persisted loss 4.3178, saved adapter, reload, and GPU inference. This proves the pipeline, not production model quality. |
| App Studio | Provides governed application projects, persistent workspaces, proposals, sessions, Git history, and isolated execution sandboxes. | Certified on Kubernetes for project access plus sandbox start, command execution, stop, persistence, and zero residual pods. |
Integrated developer and data tools
| Tool | What it does | Current verified result |
|---|---|---|
| Open WebUI | Full-featured local chat interface for installed Ollama models, files, conversations, and knowledge features. | Kubernetes-hosted, Nexus/Keycloak SSO enforced, legacy login disabled, data migrated, and real inference certified. |
| JupyterLab | Interactive Python notebooks for analysis, prototyping, data preparation, and AI experimentation. | Kubernetes-hosted; notebook creation, execution, saving, persistent notebooks, and portal rendering certified. |
| ChromaDB | Stores document embeddings and supports semantic retrieval for RAG workflows. | Kubernetes-hosted with persistent data; collection ingestion and retrieval certified. |
| Ollama | Runs local language, vision, and embedding models on the GPU host and exposes inference APIs to Nexus services. | Five installed models visible; local generation, chat, embedding-dependent workflows, and Open WebUI integration certified. |
| MLflow | Tracks experiments, parameters, metrics, and artifacts for model-development work. | Reachable through Nexus authentication; broader lifecycle migration and curated demo experiments remain a next phase. |
| LangChain | Developer framework for composing retrieval, model, tool, and application workflows from Python. | Available through the notebook/developer environment; it is a library rather than a separate end-user service. |
| FastAPI | Python API framework used by Nexus backend and custom application services. | Platform backend is deployed and serving authenticated APIs; the tool page is developer guidance. |
| NVIDIA DCGM / nvtop | Exposes GPU health, temperature, memory, utilization, and process information for operators. | GPU, driver, CUDA, and host metrics are visible through the host operations bridge and monitoring stack. |
Platform, governance, and operations
| Capability | What it does | Current verified result |
|---|---|---|
| Nexus identity and access | Central authentication, roles, sessions, and SSO for Nexus and integrated applications. | Nexus login, role enforcement, Keycloak SSO, Open WebUI SSO, Grafana SSO, and unified session use certified. |
| User and team management | Manages users, roles, teams, provisioning status, and access boundaries. | Administrative APIs and visible role enforcement are working; enterprise LDAP/AD federation is prepared but customer-specific. |
| Model catalog and router | Lists installed models and selects a model according to task policy and configured defaults. | Installed-model discovery and multiple real inference paths certified; routing policies require workload-specific tuning. |
| Guardrails, API keys, and audit | Controls programmatic access, records administrative and AI activity, and provides governance review surfaces. | Authentication, API key, audit, and role controls are integrated; formal security certification is a later phase. |
| Platform health | Reports physical host, GPU, storage, model runtime, portal, Kubernetes, Harbor, and integrated-service reachability. | Host-aware health is certified; displayed services use their actual secured or Kubernetes paths. |
| Kubernetes / K3s | Runs the portable Nexus application stack and isolated App Studio sandboxes. | Single-node K3s is healthy; scheduling, DNS, networking, storage, NVIDIA runtime, and sandbox lifecycle certified. |
| Harbor | Stores private application images and Helm charts with TLS and Nexus OIDC. | Image push/pull/delete and private OCI chart workflows certified. |
| Flux and Helm | Reconcile versioned Nexus releases from Harbor into Kubernetes. | Multiple immutable backend and portal upgrades reconciled successfully. |
| Velero and MinIO | Back up Kubernetes objects and persistent data and support recovery. | Namespace restore drill certified; final filesystem-aware backup nexus-demo-training-certified-20260721 completed 336/336 objects. |
| Prometheus and Grafana | Collect and visualize platform, host, service, and GPU metrics. | Reachable through Nexus authentication with live targets; polished product dashboards and alerts remain a next phase. |
Demonstration boundaries
- The certified LoRA job is a pipeline proof using five examples and one epoch. It is not evidence of production model quality.
- Larger 7B/8B QLoRA jobs are constrained by the reference RTX 3050’s 8 GB VRAM, particularly while Ollama is serving models.
- Audio transcription, large-model training, production alerting, enterprise directory federation, trusted customer certificates, and formal security or performance certification need environment-specific validation.
- The portal may display staging warnings for licensing, certificate, or its legacy in-application backup indicator even when Kubernetes/Velero backups exist. Presentation claims must use the verified evidence above.
Recommended next phases
- Presentation and sales-demo hardening — curated demo data, dashboards, a scripted presenter journey, screenshots, reset scripts, and a clean demo tenant.
- Model lifecycle and quality — dataset versioning, train/validation splits, quality gates, adapter publishing/merging, model registry integration, and repeatable benchmark suites.
- Observability and operations — production Grafana dashboards, alerts, centralized logs, capacity policies, and backup/restore runbooks.
- Security and enterprise identity — trusted TLS, secret rotation, LDAP/AD, MFA, vulnerability scans, rate limits, and formal role/access testing.
- Packaging and delivery — converge the Kubernetes application package and Ubuntu/K3s appliance installer, then validate online and offline upgrade, rollback, restore, and air-gapped installation paths.
- Scale and acceptance testing — concurrent users, large document sets, longer agents/workflows, representative models, hardware profiles, and customer UAT with explicit acceptance criteria.
Current deployed evidence
- Backend:
1.0.7 - Portal:
1.0.9 - Stable core checkpoint:
nexus-demo-certification-core-20260721 - Final object checkpoint:
nexus-demo-certification-final-20260721 - Dataset/training filesystem checkpoint:
nexus-demo-training-certified-20260721