Products

Arcfra Releases Neutree 1.1 for AI Inference with Native GPU Virtualization and Unified Model Governance

Published on by Arcfra Team
Last edited on

Arcfra has released Neutree 1.1, a Model-as-a-Service platform for enterprise AI inference, adding native GPU virtualization and new model governance capabilities. The release helps enterprises improve GPU utilization and manage model usage, quotas, access control, and security audits across AI applications.

Native GPU Virtualization for More Efficient Compute Usage

Enterprise AI environments often run different types of models at the same time, including LLMs, OCR, ASR, embedding, and rerank models. Traditional GPU passthrough works well for high-performance LLM inference, but lightweight models or high-concurrency small models can leave memory and compute underused when each workload occupies a full GPU.

Neutree 1.1 adds native vGPU support on top of existing GPU passthrough and logical isolation. Users can enable GPU virtualization as needed and split GPU resources by memory and compute capacity. This allows one GPU to run multiple model instances for workloads such as OCR, speech, embedding, rerank, and multi-service concurrent inference.

For performance-sensitive LLM workloads, users can still use GPU passthrough for using full card resources. This gives teams one platform for both large-model performance assurance and small-model resource reuse.

Frame 1-1.png

Neutree also provides a global view of node-level GPU usage. Administrators can see how each node and GPU is being used, then allocate, schedule, and split resources based on actual demand. With hard-isolated resource partitioning, multiple model instances can run on the same GPU without interfering with each other, helping enterprises improve GPU utilization and reduce the cost of model service delivery at scale.

New Model Governance Capabilities

As enterprises connect multiple model services to different business systems, platform teams often face four governance challenges:

  1. Limited usage visibility: Token consumption, request distribution, and cost attribution are not tracked in one place, making resource usage hard to measure.
  2. Coarse quota control: Business, apps and API Keys cannot easily receive model resources based on actual needs, which can lead to uneven resource usage.
  3. Fragmented access management: API keys are scattered across business systems, while model access scope, rate limits, and concurrency limits lack unified control.
  4. Limited traceability: Call time, access source, model, request status, and response details are not fully recorded, making troubleshooting and security audits difficult.

Frame 2.png

Neutree 1.1 extends its model gateway with a unified governance layer for internal and external models. Teams can manage usage statistics, quota management, access control, and security audits without changing existing application calling patterns.

Frame 3-1.png

Usage Statistics

Neutree tracks token usage and request activity by API key and model. Administrators can see how different apps and models consume resources, supporting capacity planning, cost allocation, and service optimization.

Quota Management

Administrators can set token quotas for each API key. This prevents one business unit or test workload from consuming too many resources and allows teams to allocate model capacity based on priority and demand.

Access Control

Neutree supports rate limits, concurrency limits, and model access scopes for each API key. This helps enterprises enforce least-privilege access, reduce misuse, and keep shared model services stable.

Access Logs and Security Audit

Neutree records request-level details, including source, API key, model, request status, token usage, throughput, latency, and finish reason. Teams can also retain request and response details for troubleshooting, performance analysis, and security audits.

Building Enterprise AI Inference Infrastructure with Arcfra AECP & Neutree

As enterprise AI moves from pilots to production, model platforms must support deployment, inference, compute management, access governance, and observability.

Neutree 1.1 helps enterprises balance LLM performance with lightweight model resource reuse through native GPU virtualization. Its model governance capabilities bring distributed model calls into one managed entry point, making usage visible, quotas controllable, permissions manageable, and requests traceable.

Neutree is a key component of Arcfra’s agentic AI infrastructure for AI inference. It focuses on compute and model management, model governance, and high-performance inference. Together with Arcfra Enterprise Cloud Platform, it helps enterprises build production-grade, high-performance, and governable model inference infrastructure.

  • Arcfra Kubernetes Engine provides a standard runtime for model inference services, supports GPU driver management, and uses high-performance networking to keep inference instances stable and efficiently scheduled.
  • Storage services provide shared and managed storage for model files, reducing duplicate storage and distribution. They can also support high-performance, high-capacity KV cache storage and persistence.
  • Observability covers GPU hardware status, container runtime status, inference instance health, model resource usage, access logs, token usage, and request latency. This gives teams end-to-end visibility for troubleshooting, security audit, performance evaluation, and continuous resource optimization.

new.png

Foxconn, the world’s largest electronics manufacturer, was among the first customers to try Neutree. Foxconn needed to modernize distributed factory infrastructure and improve how AI models were delivered, managed, and observed across environments. By using its existing Arcfra infrastructure as a production-grade AI foundation, Foxconn adopted Neutree to streamline model delivery and unify compute and model resource management. This helped Foxconn accelerate model deployment, improve operational visibility, and build a scalable foundation for future AI and intelligent manufacturing workloads.

Notes

Neutree 1.1 is now open source on GitHub. Visit the project page to learn more about features, deployment, and documentation. You can also submit issues, star the project, or join the community.

GitHub: https://github.com/neutree-ai/neutree

For Neutree Enterprise, please contact your Arcfra sales representative or submit contact requests here: https://www.arcfra.com/contact

About Arcfra

Arcfra simplifies enterprise cloud infrastructure with a full-stack, software-defined platform built for the AI era. We deliver computing, storage, networking, security, Kubernetes, and more — all in one streamlined solution. Supporting VMs, containers, and AI workloads, Arcfra offers future-proof infrastructure trusted by enterprises across e-commerce, finance, and manufacturing. Arcfra is recognized by Gartner as a Representative Vendor in full-stack hyperconverged infrastructure. Learn more at www.arcfra.com.