Arcfra today announced Neutree 1.2, a new release that expands model support and simplifies resource planning and operations management for enterprise AI platforms.
When building AI platforms, enterprises often need to manage many types of models. Beyond large language models, they also need document parsing, OCR, and traditional machine learning models for image processing, document recognition, and other use cases. These models often require extra non-inference work, such as PDF recognition and format processing, so teams usually deploy additional virtual machines or containers alongside inference engines such as vLLM. This increases the operational effort required to prepare runtime resources and makes unified API management and model governance harder across environments.
Building on standard inference engines such as vLLM and SGLang, Neutree 1.2 introduces Flex Engine to support MinerU, PaddleOCR, and selected machine learning models. Users can now manage language, multimodal, embedding, rerank, document parsing, OCR, and machine learning models on one platform, then publish and manage service calls through Neutree’s unified model gateway.

For platform teams, this reduces duplicated platform work and improves deployment efficiency. For business teams, it provides a more consistent way to access different AI capabilities and accelerates the path from model validation to business use.
Before deploying models, enterprise users often download them from public registries such as Hugging Face, then review source, parameter size, version, and other key details. As model counts grow, checking this information one by one becomes costly. Platform teams need a faster way to understand model source, size, and key attributes so they can improve model selection, deployment, and operations decisions.
Neutree 1.2 improves visual model registry management. It supports unified management of public and private models and shows source, connection status, visibility, model count, storage usage, and update time in one interface. Users can also view each model’s parameter count, size, precision, context length, and other details before deployment.

Resource planning is critical to large model deployment speed and stability. Beyond model weights, KV-cache requirements change with context window, concurrency, model architecture, and precision. Manual estimates can lead to failed deployments from insufficient resources or wasted GPU memory from over-allocation.
With the improved model repository, Neutree 1.2 can automatically parse model structure, identify key parameters, and calculate the KV-cache required for a deployment based on the user’s context window and concurrency settings. The platform combines model weights and KV-cache in the GPU memory estimate and recommends the memory needed for deployment, helping users evaluate plans against available resources.

This helps users check resources more accurately before deployment, reducing failure risk and avoiding unnecessary GPU memory waste.
As AI platforms serve more business teams, the number of API keys grows quickly. If keys are shown only in a flat list and rely on manual names or notes, administrators can struggle to understand each key’s owner and purpose. Auditing, rate limiting, and troubleshooting also become more complex.
Neutree 1.2 introduces Project-based API key management. Users can group multiple API keys by business project, select or create a Project when creating a key, and identify ownership through the Project name and description. The platform shows each key’s workspace, status, usage, rate limit, supported models, and creation time, with search, filtering, and expandable views.

This shifts API key management from a simple key list to a business-oriented view. Platform administrators can identify calling relationships and resource usage across teams more efficiently, creating a stronger foundation for governance, access audits, and resource metering.
Through standardized model delivery, unified compute and model management, model governance, gateway management, and end-to-end observability, Neutree helps enterprises move from simply deploying models to unified governance across models, compute, and applications. Together with the container, storage, networking, security, and observability capabilities of Arcfra Enterprise Cloud Platform (AECP), enterprises can build production-grade, high-performance, and governable model inference infrastructure.

Several industry customers have started using Neutree for unified model delivery and management:
An EMS Leader
This customer needed to modernize distributed factory infrastructure and improve AI model delivery, management, and observability across environments. Using its existing Arcfra infrastructure as a production-grade AI foundation, the company tested Neutree to streamline model delivery and unify compute and model resource management. This helped accelerate model deployment, improve operational visibility, and build a scalable foundation for future AI and intelligent manufacturing workloads.
A Shipbuilding Customer
This customer already used AECP to run virtualization and container workloads on one platform and then added AI capabilities. Early open-source model service trials could not meet production requirements for stability, fast rollout, unified access, operations support, and observability. The customer introduced Neutree to bring compute management, model management, and the model gateway into one platform, enabling faster model rollout, unified access, and visual observability. Neutree now supports internal shipbuilding specification queries, translation, knowledge base scenarios, and future agent applications.
A Healthcare Customer
This customer plans to build a local, manageable, and auditable MaaS platform for knowledge base Q&A, document processing, report generation, data analysis, speech transcription, and other scenarios. The goal is to provide unified, secure, and reliable model services for hospital AI applications.
The customer uses Arcfra’s HCI core ACOS, Kubernetes Engine, File Storage, and Neutree to unify compute, storage, container runtime, and model service capabilities. Model workloads run on Kubernetes Engine, while CSI provides production-grade storage and sharing for model files to support stable and efficient model services. Neutree manages LLM, OCR, ASR, and other model types in one place and uses native GPU virtualization to let multiple models share GPU resources on demand, improving utilization of limited compute. With its model gateway, the platform also provides fine-grained governance for calls, permissions, quotas, logs, and security audits, helping the customer build a governable and traceable model service environment.
Neutree 1.2 is now open source on GitHub. Visit the project page to learn more about features, deployment, and documentation. You can also submit issues, star the project, or join the community.
GitHub: https://github.com/neutree-ai/neutree
For Neutree Enterprise, please contact your Arcfra sales representative or submit contact requests here: https://www.arcfra.com/contact
Arcfra simplifies enterprise cloud infrastructure with a full-stack, software-defined platform built for the AI era. We deliver computing, storage, networking, security, Kubernetes, and more — all in one streamlined solution. Supporting VMs, containers, and AI workloads, Arcfra offers future-proof infrastructure trusted by enterprises across e-commerce, finance, and manufacturing. Arcfra is recognized by Gartner as a Representative Vendor in full-stack hyperconverged infrastructure. Learn more at www.arcfra.com.