Skip to content

BOM · Project 1 — Design-reuse Retrieval

This section contains the spec-only bill of materials (BOM) for Project 1, the design-reuse retrieval system. It lists the hardware, the software stack, and the licensing constraints, so a potential customer can size and scope a build.


D.1 The key decision: CATIA touches each part once, at ingestion

Retrieval needs geometry only to build the index, never to answer a query. So the sizing decision is even simpler than Project 2's: CATIA (if used at all) touches each part once, to produce the views, and is never in the query loop. After ingestion, embedding, indexing and search run entirely on open-source tooling.

There are two ways to get the rendered views out, mirroring Project 2's Option 1 / Option 2:

Option 1: CATIA batch
exports screenshots

Option 2: convert to STEP once,
render open-source

CATPart library

Rendered views
+ metadata

Embed views
pretrained vision model

Vector index
+ metadata payload

Similarity search
at query time

In short: CATIA or a one-time STEP conversion produces the views; everything after that is open-source and offline.

CATPart: CATIA's native part file; STEP: neutral CAD exchange format (ISO 10303).

Option 1 — CATIA exports views Option 2 — STEP-first
CATIA at ingestion Yes, a licensed seat to batch-render No; a standalone converter makes STEP, then open-source renders it
CATIA at query time No, in both options No
Operating system Windows for the export step Linux for the whole pipeline
Main risk Windows + CATIA seat coupling for the export Conversion + render fidelity on complex parts

Because Project 2 already needs a STEP path for its geometry checks, Option 2 shares that converter (Datakit CrossManager, no CATIA seat).1


D.2 BOM — compute

The GPU need rises with the phase, which suits the "start minimal, scale on evidence" staging.

Phase What runs GPU need Notes
Phase 1, shape search vision encoder, one-time batch embedding modest: a few GB VRAM; CPU-tolerable if slower DINOv2 ViT-L ≈ 300M params, ViT-g ≈ 1.1B; weights are well under ~2 GB at half precision2
Phase 2, language layer the local LLM joins 24 GB-class, the same as Project 2 reuse Project 2's compute BOM; no second GPU needed
Phase 3, B-rep re-rank self-supervised training of a B-rep encoder a training-capable GPU, one-time see D.5; optional and evidence-gated

The key point: Phase 1 is light. Embedding a design-office library (hundreds of distinct parts, a few thousand files) is a one-time job that a single modest GPU finishes quickly, and a CPU can do it if slower runs are acceptable.2 The heavy GPU appears only when the LLM does, in Phase 2.


D.3 BOM — the vision encoder

Component Recommendation Alternate Notes
Image embedding model DINOv2 (Meta) OpenCLIP (LAION) DINOv2 weights + code are Apache-2.0, cleanly commercial3; OpenCLIP code is MIT, but its checkpoints inherit their LAION training-data terms4

One honest limitation. Vision encoders are trained on everyday photographs, and untextured grey CAD renders are a known domain gap: the 3D-vision literature reports that CAD-style shaded renders, lacking texture and realistic lighting, align less well than natural images.5 Treat the magnitude as reasoned caution, not a measured number, and check it at the pilot. This is one reason the exact-geometry Phase 3 exists as a fallback.


D.4 BOM — supporting software (all open-source)

None of these needs a licence; specifying the hardware is enough.

Component Option (+ alternate) Hardware GPU?
Vector database FAISS (MIT) / Qdrant or Chroma (Apache-2.0) CPU + a few GB RAM; index is small (GBs for a design-office library) No6
Metadata database + query tool any relational or document store (e.g. SQLite / Postgres) CPU; tiny No
View rendering (Option 2) Open CASCADE / FreeCAD / trimesh CPU No
LLM serving (Phase 2) Ollama / vLLM shares the Phase-2 GPU Yes

The metadata database is where the fetch-and-filter work lives: material, customer, family, dimensions and links to past records, queried directly (Phase 1) or by the LLM (Phase 2). Building it is mostly data cleanup, the boring but mandatory part.


D.5 BOM — optional Phase-3 B-rep encoder

Phase 3 adds exact, dimension-aware matching by embedding the B-rep solid, not a screenshot. It is optional and adopted only if Phase 1 proves too coarse. The honest licensing picture:

Option Reads Licence Verdict
UV-Net (Autodesk) B-rep (from STEP) code MIT; no general/moulding-ready weights released7 Primary: train self-supervised on the firm's own parts
UVStyle-Net (Autodesk) B-rep, style similarity CC BY-NC-SA 4.0, non-commercial8 Method is the most task-aligned, but the released code is non-commercial; reimplement the approach on UV-Net's MIT encoder rather than reuse it
Point-cloud encoders (OpenShape / ULIP-2 / Uni3D) points sampled from a mesh code permissive, but weights blocked by training-data terms (Objaverse / ShapeNet)9 Item to watch: 3D-aware and avoids STEP conversion, but the weights are not cleanly commercial

Three facts to keep in mind:

  • No moulding-ready B-rep weights exist. Every option needs self-supervised training on the firm's own parts (no labels needed, but real ML effort).
  • B-rep re-imposes the STEP conversion that the render path avoids, so Phase 3 only makes sense alongside Option 2 ingestion.
  • The evidence that 3D-aware retrieval works is real but modest: the Mechanical Components Benchmark reports best-in-class retrieval around 89–91% precision, not the higher figures sometimes quoted.10

D.6 BOM — workstation baseline

  • Phase 1: a modest GPU (or CPU) for one-time embedding; 16–32 GB system RAM; NVMe SSD with room for the model, the small index and the metadata DB.
  • Phase 2: the 24 GB-class GPU from Project 2's baseline; the same box can host both projects.
  • Network: none at runtime; fully offline / LAN.

Ingestion machine: Option 2's converter needs no CATIA seat1 and can share the Linux box. Only Option 1 (CATIA batch-render) needs a licensed Windows CATIA workstation, and only during ingestion.


D.7 Licensing constraints to confirm

Item What to confirm With whom
CATIA seat (Option 1 only) That batch screenshot/render runs consume a standard seat, and licence-server capacity Dassault / CATIA sales
STEP export (Option 1 → 2) Whether batch CATPart→STEP export needs a separate CATIA licence, not publicly confirmed; do not assume Dassault / CATIA sales
Converter licensing (Option 2) CrossManager tier (GUI file-count limit vs CLI batch) and per-seat vs floating1 Datakit
OpenCLIP checkpoint terms If OpenCLIP is chosen over DINOv2, confirm the specific checkpoint's training-data terms for commercial use4 (self-check)

D.8 Open questions & limitations

Held honestly, per the study's evidence-first stance:

  • Grey-render domain gap is documented but unquantified. How much accuracy the untextured CAD look costs is not measured here; confirm at pilot.5
  • Image embeddings are scale-blind. Phase 1 recovers size only via metadata (bounding box, volume); true dimension-aware matching is the optional Phase 3.
  • No moulding-ready B-rep weights. Phase 3 means training your own encoder.
  • Point-cloud weights are licence-blocked by their training-data terms, not their code, so they are an item to watch, not a recommendation.9
  • Library size and metadata richness are unconfirmed (no access yet). Manufacturing outcomes attached to parts (gate positions, tool changes) are hoped-for until confirmed with the customer.

D.9 References


  1. Datakit CrossManager runs "on premise, without the need of external third-party CAD licenses," reads CATIA V5 CATPart/CATProduct and writes STEP, and offers a CLI tier for headless batch conversion. https://www.datakit.com/en/cross_manager.php (accessed Aug 2026). See Project 2's BOM for the full CATIA-licensing analysis. ↩↩↩

  2. DINOv2 model sizes from the official repo table: ViT-S/14 ≈ 21M, ViT-B/14 ≈ 86M, ViT-L/14 ≈ 300M, ViT-g/14 ≈ 1.1B parameters. https://github.com/facebookresearch/dinov2 (accessed Aug 2026). At half precision the weights alone are well under ~1 GB (ViT-L) to ~2 GB (ViT-g); inference adds only per-batch activations, so batch embedding fits in a few GB of VRAM. [Inference] The GB figures are reasoned from parameter count and precision, not a measured benchmark. ↩↩

  3. DINOv2 (Meta AI Research) code and pretrained weights are released under Apache-2.0. https://github.com/facebookresearch/dinov2/blob/main/LICENSE (accessed Aug 2026). ↩

  4. OpenCLIP (mlfoundations/open_clip) code is MIT; pretrained checkpoints are trained on LAION/DataComp datasets and inherit those data terms, a separate matter from the code licence. https://github.com/mlfoundations/open_clip/blob/main/LICENSE (accessed Aug 2026). ↩↩

  5. The 3D-vision literature documents a domain gap between natural photographs and untextured CAD renders: e.g. BlendCLIP notes CAD renders "fall short of providing the visual diversity and realism necessary for robust multimodal alignment," with CAD models "often featuring simplistic or entirely absent textures" (https://arxiv.org/html/2510.18244); OpenDlign motivates realistic depth-aligned renders for the same reason (https://arxiv.org/html/2404.16538v2) (both accessed Aug 2026). These support the direction qualitatively; no clean "grey vs textured, X% drop" benchmark was found. ↩↩

  6. FAISS (Meta, MIT, https://github.com/facebookresearch/faiss/blob/main/LICENSE), Qdrant (Apache-2.0, https://github.com/qdrant/qdrant), Chroma (Apache-2.0, https://github.com/chroma-core/chroma); all run fully self-hosted with no external calls (accessed Aug 2026). ↩

  7. UV-Net (Autodesk AI Lab, CVPR 2021), code MIT; the repo ships training recipes but no general-purpose pretrained weights, only narrow demo datasets (synthetic fonts for classification, machining-feature sets for segmentation), none moulding-ready. Paper https://arxiv.org/abs/2006.10211; repo https://github.com/AutodeskAILab/UV-Net (accessed Aug 2026). ↩

  8. UVStyle-Net (Autodesk AI Lab, ICCV 2021) performs unsupervised style-similarity search over B-rep CAD, built on a UV-Net encoder, but its repository is licensed CC BY-NC-SA 4.0 (non-commercial), unlike UV-Net's MIT. Paper https://arxiv.org/abs/2105.02961; repo https://github.com/AutodeskAILab/UVStyle-Net (accessed Aug 2026). ↩

  9. OpenShape (code MIT, https://github.com/Colin97/OpenShape_code), ULIP/ULIP-2 (Salesforce, code BSD-3, https://github.com/salesforce/ULIP) and Uni3D (BAAI, code MIT, https://github.com/baaivision/Uni3D) release permissive code, but their pretrained weights inherit the terms of their training data (Objaverse, ShapeNet), none of which grants a clean commercial licence for the weights (accessed Aug 2026). ↩↩

  10. Kim et al., "A Large-scale Annotated Mechanical Components Benchmark for Classification and Retrieval Tasks with Deep Neural Networks," ECCV 2020. Best reported retrieval on the benchmark is around 89–91% precision (P@N), e.g. PointCNN P@N 0.892 on Dataset A, DLAN P@N 0.912 on Dataset B. https://link.springer.com/chapter/10.1007/978-3-030-58523-5_11 (accessed Aug 2026). ↩