aipackage/docs/CURRENT_CAPABILITIES.md

11 KiB
Raw Permalink Blame History

Nexus One AI — Current Capability Catalog

Last verified: 21 July 2026
Reference environment: 172.16.10.144
Portal: https://172.16.10.144:9445
Open WebUI: https://172.16.10.144:9446

This document is the presentation and product-description source of truth for what each visible Nexus One AI tool does. “Certified” means its core action was completed against the deployed backend with a real persisted result. It does not mean every scale, hardware, security, or production-hardening scenario has been certified.

Product in one sentence

Nexus One AI is a customer-controlled enterprise AI platform that combines local model inference, governed knowledge and automation workflows, evaluation, application development, Kubernetes operations, monitoring, and recovery behind one identity and portal.

User-facing AI tools

Tool What it does Current verified result
AI Workspace Provides persistent multi-conversation chat with installed local models, system prompts, temperature controls, and image input where supported. Certified with real local inference; failed model calls are shown as failures and do not fall back to simulated answers.
Nexus Notebook / Knowledge Base Uploads and indexes documents, searches ChromaDB, and answers questions using retrieved source passages. Certified with document ingestion, vector retrieval, grounded answers, and citations.
Meeting Assistant Converts a transcript into a summary, topics, decisions, sentiment, action items, and formatted minutes. Transcript-to-minutes flow certified with real model output. Audio transcription remains dependent on supported audio and speech-model availability.
Prompt Library Stores reusable prompts by business category, tags, and document requirements. Certified for authenticated listing and administrative create, update, and delete operations.
Prompt Studio Builds prompts with personas, system instructions, parameters, optional document context, history, and hand-off to comparison or evaluation. Certified with a seeded executive demo prompt and real model execution.
Response Comparison Sends the same system and user prompt to two to four installed models and displays response, latency, and token usage side by side. Certified with genuine phi3:mini and llama3.1:8b responses. Browser history is presentation convenience, not a server-side evaluation record.
Agent Studio Defines governed multi-step agents that call models and platform tools. Certified with a completed real agent pipeline and persisted execution result.
Workflow Designer Executes ordered AI, retrieval, transformation, and integration steps. Certified with a real backend AI step; backend failure is surfaced and never replaced by simulated completion.
Scheduled Jobs Runs approved agent or workflow actions on a schedule or through a manual “Run now” action and stores the result. Certified with a real manually triggered scheduled job.
AI Evaluation Suite Runs repeatable test cases through selected models and records model and judge scores for quality, relevance, and safety. Certified with actual generation and persisted evaluation scores.
RAG Quality Measures retrieval relevance and grounded-answer quality against a selected knowledge collection. Core page and collection integration are available; quality depends on the indexed source material and evaluation set.
Feedback Review Captures user ratings and comments, supports administrative review, and promotes useful examples into evaluation workflows. Certified through feedback review and evaluation integration.
Chat Rooms Provides shared rooms for human discussion with @AI participation using local models. Certified with authenticated room messaging and a real @AI response.
Dataset Creation Converts PDF, DOCX, TXT, Markdown, or CSV sources into grounded supervised examples, supports review, and publishes approved rows as JSONL. Certified end to end: five grounded examples generated with llama3.1:8b, reviewed, approved, and published. Jobs now fail honestly if the requested valid sample count is not reached.
Model Training Fine-tunes a supported Hugging Face causal language model with a validated dataset and saves a PEFT LoRA adapter, tokenizer, configuration, logs, loss, and output path. Certified on CUDA with HuggingFaceTB/SmolLM2-135M-Instruct, five examples, one epoch, persisted loss 4.3178, saved adapter, reload, and GPU inference. This proves the pipeline, not production model quality.
App Studio Provides governed application projects, persistent workspaces, proposals, sessions, Git history, and isolated execution sandboxes. Certified on Kubernetes for project access plus sandbox start, command execution, stop, persistence, and zero residual pods.

Integrated developer and data tools

Tool What it does Current verified result
Open WebUI Full-featured local chat interface for installed Ollama models, files, conversations, and knowledge features. Kubernetes-hosted, Nexus/Keycloak SSO enforced, legacy login disabled, data migrated, and real inference certified.
JupyterLab Interactive Python notebooks for analysis, prototyping, data preparation, and AI experimentation. Kubernetes-hosted; notebook creation, execution, saving, persistent notebooks, and portal rendering certified.
ChromaDB Stores document embeddings and supports semantic retrieval for RAG workflows. Kubernetes-hosted with persistent data; collection ingestion and retrieval certified.
Ollama Runs local language, vision, and embedding models on the GPU host and exposes inference APIs to Nexus services. Five installed models visible; local generation, chat, embedding-dependent workflows, and Open WebUI integration certified.
MLflow Tracks experiments, parameters, metrics, and artifacts for model-development work. Reachable through Nexus authentication; broader lifecycle migration and curated demo experiments remain a next phase.
LangChain Developer framework for composing retrieval, model, tool, and application workflows from Python. Available through the notebook/developer environment; it is a library rather than a separate end-user service.
FastAPI Python API framework used by Nexus backend and custom application services. Platform backend is deployed and serving authenticated APIs; the tool page is developer guidance.
NVIDIA DCGM / nvtop Exposes GPU health, temperature, memory, utilization, and process information for operators. GPU, driver, CUDA, and host metrics are visible through the host operations bridge and monitoring stack.

Platform, governance, and operations

Capability What it does Current verified result
Nexus identity and access Central authentication, roles, sessions, and SSO for Nexus and integrated applications. Nexus login, role enforcement, Keycloak SSO, Open WebUI SSO, Grafana SSO, and unified session use certified.
User and team management Manages users, roles, teams, provisioning status, and access boundaries. Administrative APIs and visible role enforcement are working; enterprise LDAP/AD federation is prepared but customer-specific.
Model catalog and router Lists installed models and selects a model according to task policy and configured defaults. Installed-model discovery and multiple real inference paths certified; routing policies require workload-specific tuning.
Guardrails, API keys, and audit Controls programmatic access, records administrative and AI activity, and provides governance review surfaces. Authentication, API key, audit, and role controls are integrated; formal security certification is a later phase.
Platform health Reports physical host, GPU, storage, model runtime, portal, Kubernetes, Harbor, and integrated-service reachability. Host-aware health is certified; displayed services use their actual secured or Kubernetes paths.
Kubernetes / K3s Runs the portable Nexus application stack and isolated App Studio sandboxes. Single-node K3s is healthy; scheduling, DNS, networking, storage, NVIDIA runtime, and sandbox lifecycle certified.
Harbor Stores private application images and Helm charts with TLS and Nexus OIDC. Image push/pull/delete and private OCI chart workflows certified.
Flux and Helm Reconcile versioned Nexus releases from Harbor into Kubernetes. Multiple immutable backend and portal upgrades reconciled successfully.
Velero and MinIO Back up Kubernetes objects and persistent data and support recovery. Namespace restore drill certified; final filesystem-aware backup nexus-demo-training-certified-20260721 completed 336/336 objects.
Prometheus and Grafana Collect and visualize platform, host, service, and GPU metrics. Reachable through Nexus authentication with live targets; polished product dashboards and alerts remain a next phase.

Demonstration boundaries

  • The certified LoRA job is a pipeline proof using five examples and one epoch. It is not evidence of production model quality.
  • Larger 7B/8B QLoRA jobs are constrained by the reference RTX 3050s 8 GB VRAM, particularly while Ollama is serving models.
  • Audio transcription, large-model training, production alerting, enterprise directory federation, trusted customer certificates, and formal security or performance certification need environment-specific validation.
  • The portal may display staging warnings for licensing, certificate, or its legacy in-application backup indicator even when Kubernetes/Velero backups exist. Presentation claims must use the verified evidence above.
  1. Presentation and sales-demo hardening — curated demo data, dashboards, a scripted presenter journey, screenshots, reset scripts, and a clean demo tenant.
  2. Model lifecycle and quality — dataset versioning, train/validation splits, quality gates, adapter publishing/merging, model registry integration, and repeatable benchmark suites.
  3. Observability and operations — production Grafana dashboards, alerts, centralized logs, capacity policies, and backup/restore runbooks.
  4. Security and enterprise identity — trusted TLS, secret rotation, LDAP/AD, MFA, vulnerability scans, rate limits, and formal role/access testing.
  5. Packaging and delivery — converge the Kubernetes application package and Ubuntu/K3s appliance installer, then validate online and offline upgrade, rollback, restore, and air-gapped installation paths.
  6. Scale and acceptance testing — concurrent users, large document sets, longer agents/workflows, representative models, hardware profiles, and customer UAT with explicit acceptance criteria.

Current deployed evidence

  • Backend: 1.0.7
  • Portal: 1.0.9
  • Stable core checkpoint: nexus-demo-certification-core-20260721
  • Final object checkpoint: nexus-demo-certification-final-20260721
  • Dataset/training filesystem checkpoint: nexus-demo-training-certified-20260721