cd /news/artificial-intelligence/how-onestruction-built-the-ishigaki-… · home topics artificial-intelligence article
[ARTICLE · art-92293] src=aws.amazon.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC

ONESTRUCTION, Inc., with technical advisory from the AWS Generative AI Innovation Center (GenAIIC), built Ishigaki-IDS, a foundation model specialized for construction industry BIM workflows, to lower the barrier for non-specialists to author and manage IDS files. The model addresses data scarcity, IFC vocabulary injection, and IDS-specific grammar through a three-stage training pipeline (CPT, SFT, RLVR) and synthetic data generation, trained on Amazon EC2 P5en instances with AWS ParallelCluster. This case study, part of GENIAC Phase 3, demonstrates a reusable pattern for building domain-specialized AI models in data-scarce fields.

read7 min views1 publishedAug 11, 2026
How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC
Image: AWS ML Blog

Artificial Intelligence This post was co-written by ONESTRUCTION, Inc. and Amazon Web Services Japan G.K. as part of GENIAC (Generative AI Accelerator Challenge) Phase 3, with technical advisory from the AWS Generative AI Innovation Center (GenAIIC).

Building domain-specialized foundation models in data-scarce fields is hard. You need enough training data, specialized knowledge, and ways to verify your outputs.

ONESTRUCTION, Inc. is a construction technology startup that solves industry problems through openBIM. With technical advisory from GenAIIC, the company built Ishigaki-IDS, a foundation model (FM) specialized for construction industry BIM (Building Information Modeling) workflows. BIM is a digital representation of a building’s physical and functional characteristics, used across the construction lifecycle.

Japan’s construction sector faces a persistent labor shortage. BIM is promoted at the national level because it lets design, construction, and maintenance teams share information in one place. But adopting BIM requires specialist knowledge, and that learning cost has slowed wider use. A good example is IDS (Information Delivery Specifications), an XML-based standard that defines the information attached to and validated against a BIM model (an IFC (Industry Foundation Classes) model). Authoring an IDS file takes fluency in its grammar plus knowledge of IFC and its rules. Ishigaki-IDS lowers that barrier so practitioners who aren’t BIM specialists can review and manage attribute information.

This post is an architectural case study of how ONESTRUCTION built Ishigaki-IDS. If you’re a machine learning (ML) engineer working on domain adaptation, or a technical leader weighing how to build specialized AI models where data is scarce, you will find a pattern you can reuse. Construction and BIM professionals will also see what AI can do in their field. Familiarity with foundation model training (pre-training and fine-tuning) and basic AWS compute concepts helps, but it isn’t required.

You will learn:

  • How to use synthetic data generation to overcome data scarcity in niche domains.
  • How to build a three-stage training pipeline (CPT, SFT, RLVR) for domain specialization.
  • How to use verifiable rewards for structured output generation.
  • How to run distributed training on

Amazon Elastic Compute Cloud (Amazon EC2)P5en instanceswithAWS ParallelCluster.

Three challenges in building an IDS foundation model #

Three problems stood between us and a working IDS model.

The first was data scarcity. IDS is a relatively new standard, published in 2024, and construction in general is a domain with limited public web content. Many other domains such as finance, healthcare, and law train models on corpora of billions to hundreds of billions of tokens, but no comparable public dataset exists for IDS. Even after collecting recent web data, the volume was small and the depth was shallow, which meant the model couldn’t pick up enough context about IDS and related topics from data alone.

The second was injecting an IFC vocabulary of several thousand terms. For example, “beam” maps to IfcBeam

and “air conditioner” maps to IfcUnitaryEquipment

. This mapping has historically been done by hand by domain experts, and we needed the model to learn it directly.

The third was IDS-specific grammar. IDS is more than plain XML: its tag structure changes depending on what information is being attached or validated, and authors must use repeated patterns and dedicated tags. General-purpose foundation models struggle to produce this structure with accuracy.

Solution #

Our approach combined three ingredients: a multi-stage training pipeline, close collaboration with domain experts, and infrastructure built for stable distributed training. We start with the training pipeline.

Three-stage training pipeline

We built Ishigaki-IDS on top of Qwen3 (8B / 14B / 32B), an open-source large language model (LLM) from Alibaba Cloud known for strong multilingual capabilities and a wide range of parameter sizes. With the size range, we can experiment at smaller scales before committing to full training runs at 32B. We applied a three-stage training pipeline.

First, in continued pre-training (CPT), we injected IDS and IFC domain knowledge using web corpora plus synthetic data created with our internal domain experts. We generated valid IDS files at scale and built synthetic datasets that explained IDS-related documents from multiple angles, with synthetic data covering most of the training corpus.

Second, in supervised fine-tuning (SFT), we trained the model on pairs of IDS authoring instructions (in CSV or natural language) and their expected IDS output. SFT alone left expected issues, such as plausible but incorrect XML tag choices and wrong attribute values, so we designed a third stage to address them.

Third, in reinforcement learning with verifiable rewards (RLVR), we used IDS-Audit-Tool from buildingSMART, the international standards body, as the reward function. The tool checks XML well-formedness, IDS structural validity, and semantic consistency, so the model can iterate against mechanical correctness signals. RLVR fits the IDS task well because it refines output quality without large amounts of supervised data—useful for a data-poor domain.

Technical advisory from GenAIIC

We led development with our construction and BIM domain expertise and met with GenAIIC every two weeks for technical advisory. At each milestone, we brought training results and evaluation data to these sessions, and together we worked through five key areas:

Training data design– synthetic data strategies for the IDS domain and how to balance the data mix across CPT, SFT, and RLVR stages.** Evaluation benchmarks**– metrics covering IFC and IDS knowledge, structured generation, and general dialogue ability.** Training stages and techniques**– refining CPT, SFT, and RLVR, including long-context handling, reward shaping, and structured generation.** Training infrastructure**– parallelization, throughput, and stability for distributed training.** Result diagnosis**– when issues appeared, diagnosing root causes and setting direction for the next iteration.

Iterating on “what change improves IDS generation accuracy and practicality” at each cycle helped us build a domain-specialized foundation model in a niche, data-poor area within a short timeline.

Architecture

For the training infrastructure, we used Amazon EC2 P5en instances (two p5en.48xlarge nodes with NVIDIA H200 Tensor Core GPUs), orchestrated with AWS ParallelCluster. ParallelCluster is an open source tool that simplifies deploying and managing High Performance Computing (HPC) clusters on AWS. We stored training data, synthetic data, and checkpoints on Amazon FSx for Lustre, a fully managed file system optimized for compute-intensive workloads that delivers sub-millisecond latencies and high throughput. This setup gave us stable multi-node distributed training and parallel access to large datasets.

Evaluation

We built our own evaluation benchmark, IDS-Bench, with our internal IDS specialists. IDS-Bench measures performance across IFC version, construction discipline (architecture, structure, MEP, and common), language (Japanese and English), and the Implement, Structure, and Content axes, so the scores reflect what the model needs to handle in real work.

Results #

In our IDS-Bench evaluation, Ishigaki-IDS scored close to 100 percent on XML structural compliance and IDS structural compliance, and above 80 percent on IDS content consistency. General frontier models told a different story: they produced well-formed XML but scored under roughly 25 percent on IDS structural compliance and near 0 percent on IDS content consistency. IDS is a specialized and relatively new area, which is the kind of problem a domain-specialized model can solve. The model also supports context-length scaling with YaRN (Yet another RoPE extensioN). YaRN extends the context window of transformer models beyond their original training length without major performance degradation. We confirmed that the model generates correctly with inputs and outputs up to roughly 120k tokens.

In a joint proof-of-concept with buildingSMART, both IDS specialists and non-specialists responded positively to using the model in their work and to its ability to produce the intended IDS even from ambiguous prompts. They also gave us a list of suggestions for further development, which reinforced our view that the model is useful in practice.

Lessons learned #

Three takeaways from this project:

Synthetic data quality matters more than quantity. Our domain experts’ involvement in synthetic data creation was the difference-maker for model performance. Volume alone wouldn’t have produced the same result.Verifiable rewards accelerate iteration. UsingIDS-Audit-Tool

as an automated reward signal let us iterate faster than manual evaluation would allow, especially in a data-poor setting.Stable infrastructure lets us experiment freely. Reliable distributed training on Amazon EC2 P5en, AWS ParallelCluster, and Amazon FSx for Lustre freed us to focus on model improvements rather than debugging cluster issues.

Conclusion #

Combining domain expert collaboration, synthetic data, and RLVR tied to a verification tool worked well for building a domain-specialized model in a data-poor specialty area. Continuous technical advisory from GenAIIC helped us reach the accuracy targets measured on IDS-Bench within the GENIAC Phase 3 timeline. ONESTRUCTION will continue working with AWS to bring AI tools to the construction industry.

Next steps

If you’re interested in building domain-specialized foundation models for your industry, the following resources are a good place to start: Explore AWS GenAIIC– Learn how theAWS Generative AI Innovation Centersupports generative AI projects with technical advisory and best practices.Get started with distributed training– See theAWS ParallelCluster User Guideto set up similar infrastructure.Try Ishigaki-IDS– Access the model onHugging Faceand run it against your own IDS scenarios.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @onestruction, inc. 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-onestruction-bui…] indexed:0 read:7min 2026-08-11 ·