# VIDRAFT's AX-RAY AI Safety Diagnostic Platform Earns Alibaba ModelScope "Research Institution" Certification — Here's What Engineers Need to Know

> Source: <https://dev.to/ai_openfree_b23025ef075cf/vidrafts-ax-ray-ai-safety-diagnostic-platform-earns-alibaba-modelscope-research-institution-3i0c>
> Published: 2026-10-02 07:01:22+00:00

**TL;DR:** VIDRAFT, a Korean Pre-AGI AI startup, has become the first Korean AI organization to receive official "Research Institution" certification on Alibaba's ModelScope platform, joining a short list that includes Qwen, Shanghai AI Lab, and Z.ai (Zhipu AI). The company is now pushing AX-RAY — its structured AI safety diagnostic framework covering 12 categories and 117 risk items — into the Chinese AI ecosystem. If you're building or auditing LLMs for regulatory compliance, AX-RAY's public evaluation reports on Hugging Face are worth your attention.

**ModelScope** is a large-scale open-source AI model community co-founded by Alibaba and the China Computer Federation (CCF). It hosts over 140,000 models and serves more than 20 million users, making it China's largest AI distribution platform. While major organizations including Google, Meta, and DeepSeek maintain presences on the platform, none of them currently hold the formal "Research Institution" certification badge — a distinction VIDRAFT has now secured roughly seven months after opening its ModelScope organization page (`FINAL-Bench`) in February 2026.

**AX-RAY** is VIDRAFT's AI safety diagnostic platform. It evaluates AI models across:

Models receive a letter grade from **A through F** based on their evaluated risk profile. Diagnostic results and detailed reports are published publicly on Hugging Face.

VIDRAFT's ModelScope organization currently hosts **91 public models and 9 datasets**, including:

AX-RAY operates as a structured audit framework rather than a single benchmark task. At a conceptual level:

**Risk taxonomy construction:** The 117 risk items are derived from a cross-mapping of nine international regulatory and standards frameworks, enabling the same diagnostic run to generate compliance-relevant documentation for multiple jurisdictions simultaneously.

**Multi-category evaluation:** Each model under test is probed across all 12 safety categories. This appears to cover areas including AI agent behavior, harmful content generation, and small-model-specific failure modes — based on the public diagnostic results VIDRAFT has shared.

**Graded output:** The A–F grading system produces a human-readable safety score that can be attached to regulatory filing documentation, making it practically useful for teams preparing AI Act conformity assessments or NIST AI RMF profiles.

VIDRAFT is currently participating in a security-specialized AI foundation model development project commissioned by South Korea's Ministry of Science and ICT (MSIT) and the National IT Industry Promotion Agency (NIPA), as part of the Naver Cloud consortium.

VIDRAFT has published results from an AX-RAY diagnostic run covering **40 publicly available AI models** (domestic and international). Key findings:

These results suggest that agent-mode safety and small-model safety are currently the weakest points in the broader open-model ecosystem, at least under AX-RAY's evaluation criteria.

On Hugging Face, VIDRAFT's 82 public models and datasets accumulated **1.01 million downloads in the most recent 30-day window** at time of reporting.

VIDRAFT's models, datasets, and AX-RAY diagnostic reports are publicly accessible on both Hugging Face and ModelScope:

**Hugging Face** — browse VIDRAFT's published models and datasets (including AX-RAY evaluation reports):

```
https://huggingface.co/FINAL-Bench
```

**ModelScope** — access the certified FINAL-Bench organization page on ModelScope:

```
https://modelscope.cn/organization/FINAL-Bench
```

You can use the Hugging Face CLI to explore available assets:

```
pip install huggingface_hub
huggingface-cli scan-cache
```

No private API access, waitlist, or paid tier is mentioned for the public models and benchmark reports. The AX-RAY diagnostic *platform* (as a service for evaluating your own models) has not yet been described as publicly self-serve — check the Hugging Face organization page for current availability.

**Q: How does AX-RAY differ from existing safety benchmarks like MT-Bench or HELM?**

A: AX-RAY is explicitly designed for regulatory compliance output, not just relative model ranking. Its 117 risk items are cross-referenced against nine specific legal and standards frameworks, so a diagnostic run can directly support documentation for the EU AI Act, NIST AI RMF, and ISO/IEC 42001 in a single pass. Most academic benchmarks don't map to legal instruments in this way.

**Q: Why does ModelScope certification matter for developers outside China?**

A: ModelScope's 20M+ user base and its co-governance by the China Computer Federation make it a primary distribution channel for AI models in Chinese research and industry contexts. For teams shipping models that need to reach Chinese enterprise or academic users, having a verified "Research Institution" presence there carries similar trust signaling to a verified organization on Hugging Face. It also opens distribution to users who may not access Hugging Face directly.

**Q: Are the AX-RAY evaluation datasets themselves open?**

A: Yes — per the source, AX-RAY datasets are listed among the 9 datasets registered in VIDRAFT's ModelScope organization and are part of the 82 assets publicly available on Hugging Face.

*Originally reported by AI타임스 (2026-10-02) — [source article](https://www.aitimes.com/news/articleView.html?idxno=215897).*
