ITU-T F.748.26-2024
Technical specification for artificial intelligence cloud platforms: Performance evaluation

Standard No.
ITU-T F.748.26-2024
Release Date
2024
Published By
International Telecommunication Union (ITU)  IX  /  ITU
Latest
ITU-T F.748.26-2024
 

Introduction

Standard Overview and Technical Background

Recommendation ITU-T F.748.26, released by the International Telecommunication Union (ITU) in February 2024, is the first international technical specification for the performance evaluation of artificial intelligence cloud platforms. This standard, part of the F-series Multimedia Service Classification for Non-Telephone Telecommunication Services, marks a new phase in the standardization of AI cloud computing.

With the widespread global adoption of AI cloud platforms, the lack of unified performance evaluation standards has become a bottleneck restricting the industry's development. Different vendors use varying evaluation methods, making it difficult to compare results. The release of ITU-T F.748.26 fills this gap, providing the industry with a scientific and repeatable performance evaluation framework.


Analysis of the Core Evaluation Framework

Chapter 6 of the standard defines the overall framework for the performance evaluation of AI cloud platforms, based on three core principles:

Evaluation principles Technical requirements Implementation significance
Reproducibility Environmental configuration, workload, and evaluation indicators must be clearly defined and avoid ambiguity Ensure that different testing organizations can reproduce the evaluation results on demand
Scalability Ability to test the platform's ability to handle heavy loads and large-scale users Evaluate the platform's resilience in actual business scenarios
Practicality Applicable to mainstream AI cloud platforms, providing performance bottleneck detection and optimization guidance Provides actionable performance improvement suggestions for actual operations

The evaluation workflow consists of five key steps: configuration specification, workload input, indicator measurement, result output, and report generation. This structured design ensures the systematic and complete evaluation process.


Detailed Configuration Specification Requirements

Chapter 7 specifies the configuration information that must be clearly defined for performance evaluation, covering four dimensions:

Computing resource cluster configuration

Including node type and quantity (compute nodes, storage nodes, management nodes), topology, network protocol, and bandwidth. The topology determines the connection method and data transmission path between nodes, which directly affects the efficiency of distributed training.

Node Configuration Specifications

Detailed configuration requirements for each node:

  • Computing Devices: Type and quantity of heterogeneous computing devices such as CPUs, GPUs, and NPUs
  • Computing Power: Specific parameters such as microarchitecture, number of cores, number of threads, clock frequency, and cache
  • Memory: Type, bandwidth, speed, and capacity
  • Disk: Type, bandwidth, speed, and storage capacity

Software Configuration Requirements

Operating system type and version, hardware driver version, and runtime environment (bare metal server, virtual machine, container) all need to be clearly specified. Consistency of the software stack is critical to performance reproducibility.

Physical Environment Configuration

Physical environment parameters such as temperature, humidity range, power supply voltage, and power must be controlled within specified ranges to ensure optimal performance and longevity of the hardware.


Three-level evaluation workload and indicator system

Chapter 8 establishes a three-level evaluation system at the operation level, model level, and platform level, with specific workloads and performance indicators at each level.

Operation-level evaluation

The basic building blocks of AI computing are divided into three categories of operations:

Operation type Workload example Core indicators
Computational operations General matrix multiplication, convolution, pooling, self-attention, etc. Execution time, throughput, energy efficiency
Network operations Communication operations such as scatter, all-reduce, and all-to-all Execution time, bandwidth
IO operations Parallel IO, data replication, etc. IOPS, response time

Model-level evaluation

Based on the entire AI development lifecycle, this evaluation covers four key tasks:

Model training task: Evaluates the training performance of representative models such as ResNet, SSD, and BERT. Metrics include training time, convergence time, and distributed training speedup.

Algorithm development task: Focuses on the development environment startup time, reflecting the platform's developer-friendliness.

Model inference task: Measures inference latency and throughput, which directly impact user experience.

Model deployment task: Evaluates response latency and batch processing speedup, reflecting production environment performance.

Platform-level evaluation

Evaluates the platform's ability to handle multiple tasks simultaneously, including mixed scenarios of homogeneous and heterogeneous tasks. Key metrics include total execution time, maximum wait time, average wait time, and energy efficiency.


Evaluation Result Requirements and Reporting Specifications

Chapter 9 specifies the specific requirements for evaluation results:

Benchmark test reports must include four components: metadata, configuration information, workload description, and metric results. Metadata includes basic information such as the testing organization, test time, and test objects.

Benchmark test materials require submission of source code, test logs, model files, and metric calculation scripts to ensure the auditability and reproducibility of the evaluation process.


Implementation Recommendations and Best Practices

Appendix I provides detailed implementation recommendations:

Evaluation Program Implementation

It is recommended to implement a performance information collection interface, metric calculation functions, test logging functions, and maintain the immutability of the test program during testing.

Dataset Preparation

Apply necessary data preprocessing procedures, use a high-performance distributed file system for data storage, and maintain the integrity and order of the dataset.

Workload Execution

Avoid applying additional optimization methods (such as weight removal and model pruning) and keep the model architecture and hyperparameters unchanged during testing to ensure fair and comparable evaluation results.


Technological Evolution and Industry Impact

The release of ITU-T F.748.26 marks a new stage in the standardization of AI cloud computing. This standard, along with ITU-T F.748.11 (Deep Neural Network Processor Benchmark), ITU-T F.748.17 (AI Cloud Platform Technical Specification), and ITU-T F.748.18 (AI Multimedia Application Computing Capability Benchmark), constitutes a complete AI computing evaluation standard system.

The technical features of this standard include:

  • Multi-level assessment: Comprehensive coverage from the operational level to the platform level
  • Quantitative indicators: All indicators are quantifiable and measurable, eliminating subjective judgment
  • Reproducibility: Detailed configuration requirements ensure reproducible results
  • Practicality: Focuses on real-world application scenarios and provides practical guidance

The implementation of this standard will have a profound impact on the AI cloud platform industry: providing manufacturers with product optimization guidance, users with a basis for selection, and testing organizations with unified specifications, ultimately promoting technological progress and healthy development across the industry.


Compliance Implementation Guide

Organizations wishing to claim compliance with this Recommendation should focus on the following mandatory requirements:

In the configuration specification, the number of nodes in the computing resource cluster must be clearly specified, and the temperature range of the physical environment must be strictly controlled. During the evaluation process, any deviations beyond the mandatory requirements indicated by the phrase "as required" must be avoided.

The recommended implementation paths include: establishing a standardized testing environment, developing an automated testing tool chain, training a professional testing team, and establishing an evaluation mechanism for continuous improvement.

ITU-T F.748.26-2024 history

  • 2024 ITU-T F.748.26-2024 Technical specification for artificial intelligence cloud platforms: Performance evaluation
Technical specification for artificial intelligence cloud platforms: Performance evaluation



Copyright ©2026 All Rights Reserved
Update: Mon, 13 Jul 2026 16:12:23 +0000