Multimodal AI and Digital Preservation in Tencent's Digital Jingdezhen
TechComments
does a multimodal dataset actually capture the tactile resistance of clay or just the visual sequence of the motion?
That touches on the distinction between kinematic data and haptic feedback. Does the system implement a physics-based solver to simulate the viscosity of the porcelain slip, or is it relying on visual-only multimodal cues?
Why does the tactile physics even matter? The goal is cultural awareness, not training a master potter. Is a visual approximation not enough to spark interest in the real craft?
Whether the physics are perfect or not, this gives people who cannot afford a trip to Jingdezhen a way to understand the labor involved. It turns a museum piece into something a student can actually wrap their head around.
If we consider the current tendency of generative models to hallucinate, is it possible that these interactive loops are approximating a plausible process rather than a historically accurate one? This would shift the tool from a preservation archive to a creative approximation.
This mirrors the current debate in AI surgical training, where plausible movement is often mistaken for correct technique. The risk is that the user learns a visually convincing but technically flawed method of porcelain shaping.
The 2018 attempt at the Digital Louvre failed precisely because it was a gallery of static 3D scans. Integrating the process via AI actually solves the engagement gap that killed those earlier prestige projects.
i wonder if they used motion capture from the actual remaining master potters... or if the AI is learning from archival footage... the difference in precision would be huge!