Arcfra has released Neutree 1.1, a Model-as-a-Service platform for enterprise AI inference, adding native GPU virtualization and new model governance capabilities. The release helps enterprises improve GPU utilization and manage model usage, quotas, access control, and security audits across AI applications.
Enterprise AI environments often run different types of models at the same time, including LLMs, OCR, ASR, embedding, and rerank models. Traditional GPU passthrough works well for high-performance LLM inference, but lightweight models or high-concurrency small models can leave memory and compute underused when each workload occupies a full GPU.
Neutree 1.1 adds native vGPU support on top of existing GPU passthrough and logical isolation. Users can enable GPU virtualization as needed and split GPU resources by memory and compute capacity. This allows one GPU to run multiple model instances for workloads such as OCR, speech, embedding, rerank, and multi-service concurrent inference.
For performance-sensitive LLM workloads, users can still use GPU passthrough for using full card resources. This gives teams one platform for both large-model performance assurance and small-model resource reuse.

Neutree also provides a global view of node-level GPU usage. Administrators can see how each node and GPU is being used, then allocate, schedule, and split resources based on actual demand. With hard-isolated resource partitioning, multiple model instances can run on the same GPU without interfering with each other, helping enterprises improve GPU utilization and reduce the cost of model service delivery at scale.
As enterprises connect multiple model services to different business systems, platform teams often face four governance challenges:

Neutree 1.1 extends its model gateway with a unified governance layer for internal and external models. Teams can manage usage statistics, quota management, access control, and security audits without changing existing application calling patterns.

Neutree tracks token usage and request activity by API key and model. Administrators can see how different apps and models consume resources, supporting capacity planning, cost allocation, and service optimization.
Administrators can set token quotas for each API key. This prevents one business unit or test workload from consuming too many resources and allows teams to allocate model capacity based on priority and demand.
Neutree supports rate limits, concurrency limits, and model access scopes for each API key. This helps enterprises enforce least-privilege access, reduce misuse, and keep shared model services stable.
Neutree records request-level details, including source, API key, model, request status, token usage, throughput, latency, and finish reason. Teams can also retain request and response details for troubleshooting, performance analysis, and security audits.
As enterprise AI moves from pilots to production, model platforms must support deployment, inference, compute management, access governance, and observability.
Neutree 1.1 helps enterprises balance LLM performance with lightweight model resource reuse through native GPU virtualization. Its model governance capabilities bring distributed model calls into one managed entry point, making usage visible, quotas controllable, permissions manageable, and requests traceable.
Neutree is a key component of Arcfra’s agentic AI infrastructure for AI inference. It focuses on compute and model management, model governance, and high-performance inference. Together with Arcfra Enterprise Cloud Platform, it helps enterprises build production-grade, high-performance, and governable model inference infrastructure.

Foxconn, the world’s largest electronics manufacturer, was among the first customers to try Neutree. Foxconn needed to modernize distributed factory infrastructure and improve how AI models were delivered, managed, and observed across environments. By using its existing Arcfra infrastructure as a production-grade AI foundation, Foxconn adopted Neutree to streamline model delivery and unify compute and model resource management. This helped Foxconn accelerate model deployment, improve operational visibility, and build a scalable foundation for future AI and intelligent manufacturing workloads.
Notes
Neutree 1.1 is now open source on GitHub. Visit the project page to learn more about features, deployment, and documentation. You can also submit issues, star the project, or join the community.
GitHub: https://github.com/neutree-ai/neutree
For Neutree Enterprise, please contact your Arcfra sales representative or submit contact requests here: https://www.arcfra.com/contact
Arcfra simplifies enterprise cloud infrastructure with a full-stack, software-defined platform built for the AI era. We deliver computing, storage, networking, security, Kubernetes, and more — all in one streamlined solution. Supporting VMs, containers, and AI workloads, Arcfra offers future-proof infrastructure trusted by enterprises across e-commerce, finance, and manufacturing. Arcfra is recognized by Gartner as a Representative Vendor in full-stack hyperconverged infrastructure. Learn more at www.arcfra.com.