Sitelet https://github.com/untoldengine/UntoldEngine/blob/develop/docs/API/UsingSpatialInput.md
Skip to content

Latest commit

 

History

History
1102 lines (788 loc) · 45.6 KB

File metadata and controls

1102 lines (788 loc) · 45.6 KB

Spatial Input (Vision Pro)

Spatial input in Untold Engine follows a simple pipeline:

  1. visionOS emits raw spatial events.
  2. UntoldEngineXR converts each event into an XRSpatialInputSnapshot.
  3. Snapshots are queued in InputSystem.
  4. XRSpatialGestureRecognizer processes snapshots each frame.
  5. The engine publishes a single XRSpatialInputState your game reads in handleInput().

That separation keeps the system flexible: the OS-facing code stays in UntoldEngineXR, while gesture classification stays in the recognizer.

XR Input Model

Selection rays follow interactions

XRSpatialInputState.rayOriginWorld and rayDirectionWorld describe the selection ray supplied by a visionOS spatial event, in world coordinates. They are updated when the recognizer processes a primary interaction snapshot with a valid ray direction. This is an event-driven selection ray, not a continuously updated eye-gaze or head-forward ray. Looking around without interacting does not refresh these fields.

An event can arrive without a selection ray, including on .ended. In that case, the recognizer keeps the previous ray. Frames without new snapshots also retain the ray, currentPhase, and timestamp. Before any valid ray arrives, both ray fields are zero; clearing XR input resets the state. A nonzero ray or a .changed phase alone therefore does not indicate new input in the current frame. timestamp records the processed primary snapshot's time, not necessarily the last valid ray's time.

Use gesture signals to decide when to act:

Signal How to use it
spatialTapActive Handle a completed tap once. The recognizer sets it on .ended when the interaction did not become a drag, then clears it on the next input update. The ray may be retained from an earlier event in that interaction.
spatialPinchActive / spatialDragActive Drive ongoing manipulation while the gesture is active. These signals do not guarantee a newly received selection ray each frame.
spatialZoomActive / spatialRotateActive and their deltas Apply the current input update's two-hand gesture deltas; the recognizer clears these signals and deltas at the next update.
.cancelled End manipulation without treating the interaction as a tap. The retained ray is not a new selection.

For example, raycast against a detected real surface when a tap completes:

func handleInput() {
    let state = getXRSpatialInputState()
    guard state.spatialTapActive else { return }

    if let hit = pickRealSurfacePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        filter: .horizontalAny
    ) {
        Logger.log(message: "Tapped surface", vector: hit.worldPosition)
    }
}

Do not restrict tap handling to .began or .changed: a completed tap is reported on .ended. For entity selection, use pickedEntityId and the picked position/normal fields, which preserve the selection captured at the start of a completed tap. For ongoing transforms, use the manipulation lifecycle helpers to handle begin, update, end, and cancellation.

Gaze fields do not supply live tracking

gazePosition and gazeDirection are placeholders for future expansion. The XR runtime does not populate them with live eye-gaze or head-pose data; both default to zero.

On visionOS, InputSystem.shared.getGazeTarget(maxDistance:) only calculates gazePosition + normalize(gazeDirection) * maxDistance from the stored fields. It does not query tracking or raycast the scene. It returns nil for the default zero direction, non-finite position/direction values, or a non-finite/non-positive distance. Supplying values yourself makes this a target-point calculation, not a live gaze accessor.

The engine currently exposes no continuous gaze ray through these APIs. A head-forward direction would also be a different input from eye gaze.

The active camera entity is not the XR head pose

CameraSystem.shared.activeCamera identifies the scene's camera entity. Reading that entity with getPosition(entityId:) returns its authored or application-updated position. The XR runtime does not synchronize its entity transform to the live ARKit device anchor, so this is not a head-position accessor.

XR rendering instead derives per-eye view matrices from the ARKit device anchor and the compositor's eye transforms. That rendering pose is not published as a current head-pose accessor for game code. See the XR rendering lifecycle for how anchors are acquired and retained during tracking gaps.

What You Get in Game Code

From XRSpatialInputState, you can read:

  • spatialTapActive
  • spatialDragActive
  • spatialPinchActive
  • spatialPinchDragDelta
  • spatialZoomActive + spatialZoomDelta
  • spatialRotateActive + spatialRotateDeltaRadians
  • pickedEntityId

So your game logic can stay focused on behavior (select, move, rotate, scale), not event parsing.

Important Setup Step

You must enable XR event ingestion in your init:

func gameInit() {
    registerXREvents()
}

If you skip this, the callback still receives OS events, but the engine ignores them.

XR Input Configuration

Configure XR input behaviour with setInput before the scene starts:

// Spatial picking backend
setInput(.xr(.pickingBackend(.octreeGPUPreferred)))

// Two-hand rotate axis derivation
setInput(.xr(.twoHandRotateAxisMode(.dynamicSnapped)))

// Signal scene readiness
setInput(.xr(.sceneReady(true)))

Tune spatial manipulation thresholds with setSpatialManipulation:

setSpatialManipulation(.intentTranslationThreshold(0.01))
setSpatialManipulation(.intentRotationThreshold(0.08))
setSpatialManipulation(.classificationFrames(3))
setSpatialManipulation(.rotationSmoothing(factor: 0.25, deadzone: 0.002))
setSpatialManipulation(.zoomScale(min: 0.05, max: 20.0))

Typical Frame Usage

In your handleInput():

  • Poll getXRSpatialInputState() to get the current frame's input.
  • React to edge-triggered gestures like tap.
  • Apply continuous updates for drag/zoom/rotate while active.

For object manipulation, use the spatial manipulation free functions for robust pinch-driven transforms, then layer custom behaviour on top when needed.

Quick Example

This example shows how to drag and rotate a mesh using the engine:

func handleInput() {
    if gameMode == false { return }

    let state = getXRSpatialInputState()

    if state.spatialTapActive, let entityId = state.pickedEntityId {
        Logger.log(message: "Tapped entity: \(entityId)")
    }

    // Handles drag-based translate + twist rotation on picked entity
    processPinchTransformLifecycle(from: state)
}

What This Does

  • Tap → selects entity (via raycast picking)
  • Pinch + Drag → translates entity in world space
  • Pinch + Twist → rotates entity around a computed axis

processPinchTransformLifecycle handles:

  • Begin
  • Update
  • End
  • Cancel

This lifecycle model prevents stuck manipulation sessions.


Manipulate Parent Instead Of Picked Child

If ray picking hits a child mesh and you want to manipulate the parent actor:

var state = getXRSpatialInputState()

if let picked = state.pickedEntityId,
   let parent = getEntityParent(entityId: picked) {
    state.pickedEntityId = parent
}

processPinchTransformLifecycle(from: state)

This is useful when:

  • A character has multiple meshes
  • A building has sub-meshes
  • You want to move the root actor instead of individual geometry pieces

Important Note

Do not early-return only because pickedEntityId == nil before calling lifecycle processing.

End/cancel phases must still propagate to properly close manipulation sessions. Failing to do so can leave the engine in an inconsistent transform state.


Picking Participation And Hit Representation

Use these APIs to control whether an entity can be selected by spatial tap/ray picking and what hit representation it uses.

setEntityPickParticipation(entityId: entityId, enabled: false) // visible, not pickable
setEntityPickHitRepresentationMode(entityId: entityId, mode: .bounds) // pick using bounds
setEntityPickHitRepresentationMode(entityId: entityId, mode: .mesh) // pick using mesh (default)

Available APIs:

  • setEntityPickParticipation(entityId:enabled:)
  • getEntityPickParticipation(entityId:)
  • setEntityPickHitRepresentationMode(entityId:mode:)
  • getEntityPickHitRepresentationMode(entityId:)

Hit representation modes:

  • .none — Never pickable.
  • .bounds — Pick using bounds intersection.
  • .mesh — Pick using mesh-capable path (default behavior).

Behavior rules:

  • Default for existing entities: pick participation is enabled, hit mode is .mesh.
  • enabled == false means the entity is never returned by picking, regardless of mode.
  • mode == .none also means the entity is never returned by picking.
  • CPU and octree/GPU-preferred backends both respect these settings.

Raw Gesture Examples

It is strongly recommended to use the spatial free functions instead of raw gesture access.

Raw access is useful when:

  • You want custom manipulation behavior
  • You are building a custom editor
  • You want non-standard gesture responses

Tap (Selection)

Vision Pro air-tap gesture.

let state = getXRSpatialInputState()
if state.spatialTapActive, let entityId = state.pickedEntityId {
    // selectEntity(entityId)
}

Use this to:

  • Select objects
  • Trigger UI
  • Activate gameplay logic

Pinch Active

Single-hand pinch detected.

if InputSystem.shared.hasSpatialPinch() {
    // pinch is active
}

This does not imply dragging yet — only that a pinch is currently held.


Pinch Position

World-space position of pinch.

if let pinchPosition = InputSystem.shared.getPinchPosition() {
    // use pinchPosition
}

Useful for:

  • Placing objects
  • Spawning actors
  • Visual debugging

Pinch Drag Delta

Drag delta while pinch is active.

let state = getXRSpatialInputState()
if state.spatialPinchActive {
    let dragDelta = InputSystem.shared.getPinchDragDelta()
    // app-defined translation/scaling response
}

Common use cases:

  • Translate object along plane
  • Move UI panels
  • Drag actors in world space

Anchored Pinch Drag Helper

For stable translation (no per-frame delta accumulation), use the anchored lifecycle helper:

func handleInput() {
    let state = getXRSpatialInputState()

    processAnchoredPinchDragLifecycle(
        from: state,
        entityId: sceneRootEntity,
        dragPlane: .xz
    )
}

This helper:

  • Captures initial hand + entity world positions
  • Applies absolute displacement from gesture start
  • Optionally constrains world-axis movement to .xy, .xz, or .yz
  • Optionally transforms the final world position before it is written
  • Cleans up session state on end/cancel

Use this when moving large roots (buildings/scenes) where incremental delta jitter can become visible. Use .xz to preserve height while dragging across the floor axes, .xy to preserve depth for wall-style movement, and .unconstrained for free 3D movement.

dragPlane filters the hand displacement in world axes. It does not raycast the input ray against a mathematical plane. For ray-plane picking, use pickGroundPosition or pickPlanePosition.

Use positionTransform for continuous snapping, clamping, or custom placement rules:

processAnchoredPinchDragLifecycle(
    from: state,
    entityId: sceneRootEntity,
    dragPlane: .xz,
    positionTransform: { worldPosition in
        let gridSize: Float = 0.25
        return simd_float3(
            (worldPosition.x / gridSize).rounded() * gridSize,
            worldPosition.y,
            (worldPosition.z / gridSize).rounded() * gridSize
        )
    }
)

The closure receives and returns world-space position after sensitivity and dragPlane have been applied. Its return value is the final position, so it can intentionally override the constrained axis. If it returns a non-finite value, the engine skips that frame's position write.


Anchored Scene Drag Helper

For translating the entire scene root (rather than a single entity), use the anchored scene drag lifecycle:

func handleInput() {
    let state = getXRSpatialInputState()

    processAnchoredSceneDragLifecycle(from: state)
}

This helper:

  • Captures initial hand + scene root world positions on drag start
  • Applies absolute displacement from gesture start via translateSceneTo, keeping static batches intact
  • Cleans up session state on end/cancel

You can adjust movement speed with the sensitivity parameter (defaults to 1.0):

processAnchoredSceneDragLifecycle(from: state, sensitivity: 0.5)

To manually end the drag (e.g. on a mode change), call:

endAnchoredSceneDrag()

Use this when panning an entire scene — for example, sliding a map, architectural model, or level layout in world space.


Anchored Scene Rotate Helper

For rotating the entire scene root around world up (+Y) while preserving static batching, use the anchored scene rotate lifecycle. This requires a two-hand pinch + twist gesture (spatialRotateActive with both hands pinching):

func handleInput() {
    let state = getXRSpatialInputState()

    processAnchoredSceneRotateLifecycle(from: state)
}

This helper:

  • Activates only when both hands are pinching and a two-hand rotate gesture is recognized
  • Captures the initial two-hand vector direction + scene yaw on rotate start
  • Applies absolute yaw from gesture start via rotateSceneToYaw, keeping static batches intact
  • Ends automatically when either hand releases or the rotate gesture ends

You can adjust rotation speed with the sensitivity parameter (defaults to 1.0):

processAnchoredSceneRotateLifecycle(from: state, sensitivity: 0.5)

To manually end rotation (e.g. on a mode change), call:

endAnchoredSceneRotate()

Use this when aligning or calibrating an already-loaded large scene in place without rebatching.


Unified Scene Manipulation Helper

To avoid drag/rotate gesture fighting, use the unified scene-root manipulation lifecycle:

func handleInput() {
    let state = getXRSpatialInputState()

    processAnchoredSceneManipulationLifecycle(
        from: state,
        dragSensitivity: 1.0,
        rotateSensitivity: 0.5
    )
}

Arbitration rules:

  • When a pinch is first detected, classification is deferred for a few frames so the second hand has time to arrive
  • Two-hand pinch + twist (spatialRotateActive + both hands pinching) routes to scene rotate
  • Otherwise, after the deferral window expires, pinch drag routes to scene drag
  • The non-winning session is ended automatically
  • Once a mode is chosen, it stays latched (drag or rotate) until the gesture ends

You can tune the deferral window:

setSpatialManipulation(.classificationFrames(4))  // ~44ms at 90 Hz

To manually end the unified lifecycle (e.g. on a mode change), call:

endAnchoredSceneManipulation()

Use this as the default scene-root helper when your app supports both panning and rotation.


Combining Scene Drag, Rotate and Zoom

Scene drag/rotate and entity zoom can share an input loop, but choose explicit precedence: the unified lifecycle latches its drag/rotate mode, and adding a second hand does not automatically stop an already active drag. This example gives two-hand entity zoom priority and ends the scene session before scaling:

func handleInput() {
    let state = getXRSpatialInputState()

    if state.currentPhase != .ended, state.currentPhase != .cancelled,
       state.leftHandPinching, state.rightHandPinching, state.spatialZoomActive {
        endAnchoredSceneManipulation()
        applyTwoHandZoomIfNeeded(from: state)
    } else {
        processAnchoredSceneManipulationLifecycle(from: state)
    }
}

applyTwoHandZoomIfNeeded changes a selected entity's local scale (its parent by default); it does not resize SceneRootTransform. For a tabletop model or an entire placed scene, use the scene-root scaling recipe below instead.

For context-based entity vs. scene rotation — route two-hand twist to entity rotate when something is picked, and to scene rotate otherwise:

func handleInput() {
    let state = getXRSpatialInputState()

    if state.currentPhase == .ended || state.currentPhase == .cancelled {
        endAnchoredSceneManipulation()
        return
    }

    if state.leftHandPinching, state.rightHandPinching, state.spatialZoomActive {
        endAnchoredSceneManipulation()
        applyTwoHandZoomIfNeeded(from: state)
        return
    }

    if state.pickedEntityId != nil {
        endAnchoredSceneManipulation()
        // Entity is picked → two-hand twist rotates the entity
        applyTwoHandRotateIfNeeded(from: state)
    } else {
        // Nothing picked → drag or rotate the scene
        processAnchoredSceneManipulationLifecycle(from: state)
    }
}

Two-Hand Zoom

Apply the built-in entity zoom response. This scales the target's local transform, even if you call that entity your "scene root"; it does not change the engine's SceneRootTransform. Its limits come from setSpatialManipulation(.zoomScale(min:max:)).

let state = getXRSpatialInputState()

applyTwoHandZoomIfNeeded(from: state, sensitivity: 1.0)

By default, the helper scales the parent of the picked entity when available. If you want to choose the exact target, pass entityId:

let state = getXRSpatialInputState()

if let picked = state.pickedEntityId {
    // Scale exactly what was hit
    applyTwoHandZoomIfNeeded(from: state, entityId: picked, sensitivity: 1.0)

    // Or scale its parent explicitly
    if let parent = getEntityParent(entityId: picked) {
        applyTwoHandZoomIfNeeded(from: state, entityId: parent, sensitivity: 1.0)
    }
}

Two-Hand Scene-Root Scaling

To resize an entire placed scene, use the public uniform scaleSceneTo(_ scale: Float) overload. It changes SceneRootTransform without rewriting entity transforms or rebuilding static batches. No picked entity is required. getSpatialZoomDelta() reports the frame's change in hand separation in meters (positive for spreading, negative for bringing hands together). The following app-defined response maps that delta through sensitivity to a multiplicative factor (1 + delta * sensitivity), once per frame, matching the entity helper's response:

import UntoldEngine
import simd

// Call scaleSceneTo(1.0) when initializing/resetting your placement.
// Configure these limits for your model's units and intended tabletop size.
func applySceneRootZoom(
    from state: XRSpatialInputState,
    sensitivity: Float = 1.0,
    minScale: Float = 0.05,
    maxScale: Float = 20.0
) {
    guard state.currentPhase != .ended, state.currentPhase != .cancelled,
          state.leftHandPinching, state.rightHandPinching, state.spatialZoomActive,
          sensitivity.isFinite, sensitivity > 0,
          minScale.isFinite, maxScale.isFinite,
          minScale > 0, maxScale >= minScale else { return }

    let delta = InputSystem.shared.getSpatialZoomDelta() * sensitivity
    guard delta.isFinite, delta != 0 else { return }

    let factor: Float = 1 + delta
    guard factor.isFinite, factor > 0 else { return }

    let current = SceneRootTransform.shared.scale
    // This recipe expects an already uniform, positive scene scale.
    guard current.x.isFinite, current.y.isFinite, current.z.isFinite,
          current.x > 0, current.x == current.y, current.x == current.z else { return }

    let requested = current.x * factor
    guard requested.isFinite else { return }

    let clamped = min(max(requested, minScale), maxScale)
    scaleSceneTo(clamped) // Float overload writes the same value on all three axes.
}

func handleInput() {
    let state = getXRSpatialInputState()

    if state.currentPhase != .ended, state.currentPhase != .cancelled,
       state.leftHandPinching, state.rightHandPinching, state.spatialZoomActive {
        // Zoom owns this frame, even if the delta is rejected or at a limit.
        endAnchoredSceneManipulation()
        applySceneRootZoom(from: state, minScale: 0.05, maxScale: 20.0)
    } else {
        // Always forward end/cancel and release frames for lifecycle cleanup.
        processAnchoredSceneManipulationLifecycle(from: state)
    }
}

These are app-configured scene limits, independent of the entity helper's .zoomScale setting. Invalid limits, sensitivity, deltas, overflow, or a nonuniform/nonpositive current scale skip the write. Initialize the scene with a uniform scale inside your chosen range; after a valid update the scale stays within that range. Poll and apply the delta only once in each handleInput() frame; do not also call the entity zoom helper for the same scene-resize interaction.

Placement Pivot And Lifecycle Composition

The scene-root matrix is T * R * S. Scaling keeps the scene-local origin fixed at SceneRootTransform.shared.position in visual world space; it changes neither position nor rotation. Author or arrange the model so its intended contact point (for example, the center of its base) is at scene-local (0, 0, 0), then place that origin on the table using translateSceneTo(position: hit.worldPosition). A contact point away from the origin will move when scaling: the pivot is not the picked entity, hand midpoint, or last surface hit.

The input loop gives zoom priority over rotation if both signals are active. Ending the unified lifecycle clears its latched mode and cached drag position/rotation baseline. When zoom ends, drag/rotate begins a fresh session at the current placement, so a drag that started before zoom cannot restore an old placement. This intentionally sequences zoom and drag/rotate; it does not provide simultaneous three-way manipulation. Call endAnchoredSceneManipulation() on mode changes and before externally changing placement. If your app previously used the individual scene drag/rotate helpers, switch to this unified loop rather than continuing to call them alongside it.

For a contact point away from the origin, keeping that point fixed would also require compensating the root translation: for scene-local pivot p, preserve worldPivot = position + rotation.act(scale * p) and set the new position to worldPivot - rotation.act(newScale * p). Such translation must end/rebase any anchored drag session. It also changes the root origin used by scene yaw rotation, so the origin-based placement above is the recommended recipe when composing with the existing anchored rotation helper.


Two-Hand Rotate

Configure how the rotation axis is derived:

setInput(.xr(.twoHandRotateAxisMode(.dynamicSnapped)))

Available modes:

  • .cameraForward — rotates around camera-forward axis (screen-style twist)
  • .dynamic — derives axis from actual two-hand motion
  • .dynamicSnapped — dynamic axis snapped to dominant world axis (x, y, or z)

Apply the built-in rotate response:

let state = getXRSpatialInputState()

applyTwoHandRotateIfNeeded(from: state, sensitivity: 1.5)

By default, the helper rotates the parent of the picked entity when available. If you want to choose the exact target, pass entityId:

let state = getXRSpatialInputState()

if let picked = state.pickedEntityId {
    // Rotate exactly what was hit
    applyTwoHandRotateIfNeeded(from: state, entityId: picked, sensitivity: 1.5)

    // Or rotate its parent explicitly
    if let parent = getEntityParent(entityId: picked) {
        applyTwoHandRotateIfNeeded(from: state, entityId: parent, sensitivity: 1.5)
    }
}

Get distance to hit-entity

To get the distance to an entity use the following:

let state = getXRSpatialInputState()
if state.spatialTapActive, let entityId = state.pickedEntityId {
    let distance = state.pickedEntityDistance
    print("Object distance: \(distance) meters")
}

Get Model Hit Position And Normal

When a spatial tap hits a pickable model/entity, XRSpatialInputState includes the exact world-space hit point. GPU mesh picking can also provide the geometric surface normal of the triangle that was hit.

let state = getXRSpatialInputState()

if state.spatialTapActive, let entityId = state.pickedEntityId {
    let modelPoint = state.pickedEntityWorldPosition
    let modelNormal = state.pickedEntityWorldNormal

    Logger.log(message: "Picked model entity: \(entityId)")

    if let modelPoint {
        Logger.log(message: "Model hit point", vector: modelPoint)
    }

    if let modelNormal {
        Logger.log(message: "Model surface normal", vector: modelNormal)
    }
}

pickedEntityWorldNormal is optional. It is populated by GPU mesh picking for triangle hits. CPU bounds picking and other non-mesh fallback paths do not report a true model surface normal, so they leave this value nil.


Get Ground/Plane Hit Position

To retrieve the exact world-space position where the user taps on a real-world surface, use pickRealSurfacePosition. This raycasts against ARKit-detected physical planes in the user's environment. This is useful for calibration workflows where you need to anchor a point on the ground and scale a model relative to it.

The filter parameter controls which planes are considered by alignment and, optionally, by surface classification. The function always returns the single closest hit that passes the filter.

Hit fields and coordinate spaces

pickRealSurfacePosition returns RealSurfaceHit?; it returns nil when no tracked plane qualifies. All input rays and returned positions use the session's physical world coordinates, not authored scene/entity coordinates.

Field Meaning
worldPosition: simd_float3 Intersection point in physical world space, in meters.
surfaceKind: RealSurfaceKind Detected plane classification: .floor, .ceiling, .wall, .table, .seat, .door, .window, or .unknown.
distance: Float Distance from the input ray origin to worldPosition, in world-space meters.
planeNormal: simd_float3 Unit normal of the detected plane in world space.
surfaceNormal: simd_float3 Read-only alias of planeNormal.

Pass XRSpatialInputState.rayOriginWorld and rayDirectionWorld directly. The direction does not need to be normalized. hitYRange tests the intersection's physical world Y coordinate in meters. maxDistance limits the accepted hit distance in world-space meters and defaults to unlimited. Both maxDistance and the returned distance remain in physical meters when SceneRootTransform scales the authored scene; do not rescale the input ray or distance limit.

Where detected planes come from

TrackedPlane is the core engine's platform-independent snapshot of a detected plane. It contains the anchor id, originFromAnchorTransform (anchor to physical world), anchorFromExtentTransform (extent to anchor), extentWidth and extentHeight in meters along the extent's local X and Z axes, alignment, and classification. Picking tests the finite extent rectangle, including its offset from the anchor origin.

RealSurfacePlaneStore.shared is the thread-safe store of the latest [TrackedPlane]. The UntoldEngineXR layer populates it from ARKit's PlaneDetectionProvider.anchorUpdates, handling added, updated, and removed anchors. The core UntoldEngine picking function reads a snapshot() of that store; it does not start plane detection itself. See the XR plane monitor for the provider lifecycle.

Normal XR application code consumes the store rather than populating it. update(planes:) replaces the entire set, snapshot() reads it, clear() empties it, and logAllPlanes() prints diagnostics. Without plane data, picking returns nil; check world-sensing authorization and allow time for scanning.

Tabletop placement with a scaled scene

This example shrinks the authored scene to 1% of its size and places a top-level entity's origin on a detected table. Call the setup once, then call the placement function from your tap/update handler with an existing entity that has a transform:

import simd
import UntoldEngine

func configureTabletopScale() {
    SceneRootTransform.shared.scale = simd_float3(repeating: 0.01)
    SceneRootTransform.shared.updateIfNeeded()
}

func placeOnTable(entityId: EntityID) {
    let state = getXRSpatialInputState()
    guard state.spatialTapActive,
          let hit = pickRealSurfacePosition(
              rayOrigin: state.rayOriginWorld,
              rayDirection: state.rayDirectionWorld,
              filter: .tableOnly,
              maxDistance: 2.0 // Two physical meters, even at 1% scene scale.
          )
    else { return }

    let scenePoint = SceneRootTransform.shared.visualWorldToSceneLocal(hit.worldPosition)
    translateTo(entityId: entityId, position: scenePoint)
}

worldPosition stays on the physical table. Convert it with visualWorldToSceneLocal(_:) before using it as an authored scene position; the helper accounts for scene-root translation, rotation, and scale. With only a uniform 0.01 scale, a world point (0, 0.75, 0) becomes scene point (0, 75, 0), while a hit one physical meter from the ray origin still reports distance == 1. Use sceneLocalToVisualWorld(_:) for the reverse conversion.

translateTo sets an entity's local position. The example assumes a top-level entity; for a child, also convert the scene point into its parent's local space before passing it to translateTo. If the table remains .unknown, use .horizontalAny with a hitYRange chosen from observed world heights, as described below.

Alignment presets

  • .horizontalAny — horizontal planes only (floor, ceiling, table, seat). Warning: this includes tables and seats — use .floorOnly when you need the floor specifically.
  • .verticalAny — vertical planes only (wall, door, window)
  • .any — all detected planes regardless of alignment

Classification presets

  • .floorOnly — floor planes only (recommended for ground anchoring)
  • .tableOnly — table planes only
  • .wallOnly — wall planes only

Picking whichever surface the user is pointing at

When your app needs to respond to floor or table (whichever the user taps), use a single call with a multi-kind filter and inspect surfaceKind in the result. Because the function returns the closest qualifying hit, this correctly returns the table when pointing at the table and the floor when pointing at the floor.

let state = getXRSpatialInputState()

if state.spatialTapActive {
    let filter = RealSurfaceFilter(alignment: .horizontal, kinds: [.floor, .table])

    if let hit = pickRealSurfacePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        filter: filter
    ) {
        switch hit.surfaceKind {
        case .floor:
            Logger.log(message: "Floor hit", vector: hit.worldPosition)
        case .table:
            Logger.log(message: "Table hit", vector: hit.worldPosition)
        default:
            break
        }
    }
}

Anti-pattern — do not call pickRealSurfacePosition twice in the same tap handler with different classification filters. Each call is an independent ray cast. When pointing at a table, a .floorOnly call will skip the table plane and keep going until it hits the large floor plane behind it — so both calls return a hit even though the user only pointed at one surface. Use a single call and branch on surfaceKind.

Other filter examples

let state = getXRSpatialInputState()

if state.spatialTapActive {
    // Floor only — always ignores tables, seats, and ceilings
    if let hit = pickRealSurfacePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        filter: .floorOnly
    ) {
        Logger.log(message: "Floor hit", vector: hit.worldPosition)
    }

    // Any horizontal surface — inspect kind after the fact
    if let hit = pickRealSurfacePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        filter: .horizontalAny
    ) {
        Logger.log(message: "Surface type: \(hit.surfaceKind)", vector: hit.worldPosition)
    }

    // Vertical surface (wall, door, window)
    if let hit = pickRealSurfacePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        filter: .verticalAny
    ) {
        Logger.log(message: "Surface type: \(hit.surfaceKind)", vector: hit.worldPosition)
    }
}

RealSurfaceHit includes both the tap point and the detected plane normal:

if let hit = pickRealSurfacePosition(
    rayOrigin: state.rayOriginWorld,
    rayDirection: state.rayDirectionWorld,
    filter: .wallOnly
) {
    let wallPoint = hit.worldPosition
    let wallNormal = hit.surfaceNormal

    Logger.log(message: "Wall point", vector: wallPoint)
    Logger.log(message: "Wall normal", vector: wallNormal)
}

Use .wallOnly when the workflow specifically needs an ARKit-classified wall. Use .verticalAny when doors, windows, or temporarily unknown vertical planes are acceptable.

Aligning a model wall to a real wall

For digital twin or floor-plan calibration, use one tap on the virtual model wall and one tap on the detected real wall. The model tap provides pickedEntityWorldPosition plus pickedEntityWorldNormal; the real wall tap provides RealSurfaceHit.worldPosition plus surfaceNormal.

struct WallAlignmentSample {
    var modelPoint: simd_float3?
    var modelNormal: simd_float3?
    var realPoint: simd_float3?
    var realNormal: simd_float3?
}

var wallAlignment = WallAlignmentSample()

func captureModelWallTap() {
    let state = getXRSpatialInputState()

    guard state.spatialTapActive,
          state.pickedEntityId != nil,
          let point = state.pickedEntityWorldPosition,
          let normal = state.pickedEntityWorldNormal
    else {
        return
    }

    wallAlignment.modelPoint = point
    wallAlignment.modelNormal = normal
}

func captureRealWallTap() {
    let state = getXRSpatialInputState()

    guard state.spatialTapActive,
          let hit = pickRealSurfacePosition(
              rayOrigin: state.rayOriginWorld,
              rayDirection: state.rayDirectionWorld,
              filter: .wallOnly
          )
    else {
        return
    }

    wallAlignment.realPoint = hit.worldPosition
    wallAlignment.realNormal = hit.surfaceNormal
}

func applyWallAlignment(to modelRoot: EntityID) {
    guard let modelPoint = wallAlignment.modelPoint,
          let modelNormal = wallAlignment.modelNormal,
          let realPoint = wallAlignment.realPoint,
          let realNormal = wallAlignment.realNormal
    else {
        return
    }

    let modelRootPosition = getPosition(entityId: modelRoot)
    let currentOrientation = getOrientation(entityId: modelRoot)
    let currentRotation = transformMatrix3nToQuaternion(m: matrix_float3x3(columns: (
        simd_float3(currentOrientation.columns.0.x, currentOrientation.columns.0.y, currentOrientation.columns.0.z),
        simd_float3(currentOrientation.columns.1.x, currentOrientation.columns.1.y, currentOrientation.columns.1.z),
        simd_float3(currentOrientation.columns.2.x, currentOrientation.columns.2.y, currentOrientation.columns.2.z)
    )))

    let deltaRotation = simd_quatf(from: simd_normalize(modelNormal), to: simd_normalize(realNormal))
    let targetRotation = simd_normalize(deltaRotation * currentRotation)
    let rotatedModelPoint = modelRootPosition + deltaRotation.act(modelPoint - modelRootPosition)
    let targetPosition = modelRootPosition + (realPoint - rotatedModelPoint)

    rotateTo(entityId: modelRoot, rotation: getMatrix4x4FromQuaternion(q: targetRotation))
    translateTo(entityId: modelRoot, position: targetPosition)
}

This sample assumes modelRoot is the top-level model entity you want to calibrate and SceneRootTransform is identity. With a scene-root transform, convert the world points and normals into authored scene space before computing the alignment. If the entity has a parent transform, convert the target position and rotation into that parent space before calling translateTo or rotateTo. In a production calibration flow, keep the model's intended up axis stable when applying the rotation so wall alignment does not introduce unwanted roll.

Choosing the right filter

Goal Filter to use
Always anchor to the floor, ignore furniture .floorOnly
Always anchor to the table, ignore floor .tableOnly
Whichever surface the user taps kinds: [.floor, .table] + check surfaceKind
Any horizontal surface .horizontalAny + check surfaceKind
Real wall normal for model alignment .wallOnly + surfaceNormal

Diagnosing unexpected classification

If surfaces are not being detected as expected, call this at any point to print every plane ARKit currently tracks, including its classification, Y position, and size:

RealSurfacePlaneStore.shared.logAllPlanes()

Sample output:

── RealSurfacePlaneStore: 3 plane(s) ──────────────────
  [a1b2c3d4] alignment=horizontal  classification=floor    y=-0.02m  size=4.20x3.80
  [e5f6a7b8] alignment=horizontal  classification=unknown  y=+0.74m  size=1.10x0.60
  [c9d0e1f2] alignment=vertical    classification=wall     y=+1.20m  size=2.40x0.10
────────────────────────────────────────────────────────────────────

This reveals a common issue: ARKit frequently classifies desks and tables as .unknown rather than .table, especially when the surface has not been scanned from multiple angles or the room lighting is poor. Waiting and walking around the furniture can help ARKit reclassify.

Targeting surfaces by height (Y-range filter)

When ARKit does not classify a desk or table correctly, use the hitYRange parameter to restrict hits by the world-space Y coordinate of the intersection point. This is reliable regardless of classification.

The floor's Y coordinate depends on the session's coordinate origin; it is not universally near Y≈0. Choose height filters from observed plane positions (use logAllPlanes()), rather than assuming a fixed floor height. The following ranges assume a session where the observed floor is near Y=0 and the desk or table is between Y=0.5m and Y=1.1m. Adjust both ranges for your session; scene-root scaling does not change these physical world heights.

let state = getXRSpatialInputState()

if state.spatialTapActive {
    // Floor — accept hits within ±20 cm of ground level
    if let hit = pickRealSurfacePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        filter: .horizontalAny,
        hitYRange: (-0.2)...0.2
    ) {
        Logger.log(message: "Floor hit (Y=\(hit.worldPosition.y))", vector: hit.worldPosition)
    }

    // Desk or table — accept hits between 0.5m and 1.1m
    if let hit = pickRealSurfacePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        filter: .horizontalAny,
        hitYRange: 0.5...1.1
    ) {
        Logger.log(message: "Desk hit (Y=\(hit.worldPosition.y))", vector: hit.worldPosition)
    }
}

You can combine hitYRange with a classification filter. When ARKit does classify surfaces correctly this gives the tightest constraint:

if let hit = pickRealSurfacePosition(
    rayOrigin: state.rayOriginWorld,
    rayDirection: state.rayDirectionWorld,
    filter: .floorOnly,
    hitYRange: (-0.2)...0.2
) { ... }

Note on ARKit classification timing

ARKit can initially report a newly-detected horizontal plane as .unknown before it has gathered enough geometry to classify it as floor or table. If placement feels unreliable immediately after startup, wait a few seconds and walk around the surface to give ARKit more data. Use logAllPlanes() to monitor classification as it updates.


Get Virtual Plane Hit Position

Virtual planes are purely mathematical — no ARKit scanning required. Use them when you want to cast a ray against a plane you define in code rather than one detected from the real environment. Common cases: snapping objects to the engine's ground level (Y = 0), placing content on a wall you defined, or constraining drag to an arbitrary surface.

Two functions are available, both returning a PlanePickHit with worldPosition and distance:

Horizontal ground plane

pickGroundPosition casts against a horizontal plane at a given Y height. planeY defaults to 0.

let state = getXRSpatialInputState()

if state.spatialTapActive {
    if let hit = pickGroundPosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        planeY: 0.0
    ) {
        Logger.log(message: "Virtual ground hit", vector: hit.worldPosition)
    }
}

Use planeY to match a raised or sunken surface — for example planeY: 0.75 for a table-height virtual plane.

Arbitrary virtual plane

pickPlanePosition casts against any plane defined by a world-space point and normal.

let state = getXRSpatialInputState()

if state.spatialTapActive {
    // Vertical plane facing +Z, passing through the origin
    if let hit = pickPlanePosition(
        rayOrigin: state.rayOriginWorld,
        rayDirection: state.rayDirectionWorld,
        planePoint: simd_float3(0, 0, 0),
        planeNormal: simd_float3(0, 0, 1)
    ) {
        Logger.log(message: "Virtual wall hit", vector: hit.worldPosition)
    }
}

planePoint can be any point that lies on the plane — the normal does not need to be pre-normalized.

Choosing between virtual and real

Goal Function to use
Snap to engine ground (Y = 0) pickGroundPosition
Snap to a raised virtual surface pickGroundPosition(planeY:)
Cast against a wall or angled surface you defined pickPlanePosition
Cast against a pickable model/entity surface pickedEntityWorldPosition + pickedEntityWorldNormal
Cast against an ARKit-detected physical surface pickRealSurfacePosition

Both pickGroundPosition and pickPlanePosition automatically account for scene root transforms, so the math stays correct even when the scene has been translated or rotated.


Spatial Helper Functions

Use these free functions for spatial manipulation. They all delegate to SpatialManipulationSystem internally so you never need to reference the shared singleton directly.

  • processPinchTransformLifecycle(from:) Recommended default. Handles translation + twist rotation lifecycle safely.

  • applyPinchDragIfNeeded(from:entityId:sensitivity:) Lower-level translation helper if you want full control.

  • processAnchoredPinchDragLifecycle(from:entityId:sensitivity:dragPlane:positionTransform:) Anchored drag for a single entity. Applies absolute displacement from gesture start, optionally constrained by world-axis displacement filtering.

  • processAnchoredSceneDragLifecycle(from:sensitivity:) Anchored drag for the entire scene root. Applies absolute displacement via translateSceneTo.

  • endAnchoredSceneDrag() Manually ends an in-progress anchored scene drag session.

  • processAnchoredSceneRotateLifecycle(from:sensitivity:) Anchored rotate for the entire scene root using two-hand pinch + twist. Applies absolute yaw via rotateSceneToYaw.

  • endAnchoredSceneRotate() Manually ends an in-progress anchored scene rotate session.

  • processAnchoredSceneManipulationLifecycle(from:dragSensitivity:rotateSensitivity:) Unified scene-root helper with drag/rotate arbitration to prevent gesture-fighting.

  • endAnchoredSceneManipulation() Ends any in-progress unified scene manipulation (drag, rotate, or pending classification).

  • applyTwoHandZoomIfNeeded(from:entityId:sensitivity:) Scales the picked entity (or its parent) using the two-hand spread/pinch gesture.

  • applyTwoHandRotateIfNeeded(from:entityId:sensitivity:axisOverrideWorld:) Rotates the picked entity (or its parent) using the two-hand twist gesture.

  • endSpatialManipulation() Ends the current pinch-transform manipulation session.

  • resetSpatialManipulation() Resets all manipulation session state (use when changing modes or reloading scenes).