Repository navigation
How to train <task-name>-real? #3
Copy link
Copy link
Open
Description
Activity
I appreciate your interest in DreamControl!
- Yes, you are right, there should be a platform during training. This is something I missed while refactoring the code; I have now resolved it. Sorry about the error. You should be able to train the policy now. Please let me know if you still face any issues
- Resolved it now. Again, this is something I missed while refactoring the code
- 2.1 is an arbitrary value chosen to get a good camera view for eval video. It could be set to any other value. The reason behind standardizing the x position of the object is that IsaacLab does not directly support moving (or respawning) non-rigidbody objects during training for each new episode. As I want the platform to be static and not a rigidbody, what I did was move the starting point of the motion to be able to fix the platform at 2.1 (or any other fixed value). Let me know if you need more explanation
- The idea is that you generate a large no of motions, randomizing the object grasp positions within a distribution, and use those as a reference to train a policy to grasp an object within that range. During test time, the policy should generalize to grasping an object located within the same range as training. Although the sampled positions may not be directly present in the set of generated trajectories used for training, the no of those trajectories used for training is large enough to force the policy to generalize to any other position within the same distribution
- The code only uses USD files from humanoidverse. I have removed all other scripts from the repo to avoid confusion. Thanks for pointing that out!
Let me know if you run into other issues. Belated Merry Christmas!
Best,
DvijThanks for your reply — it clarified a lot of things for me. I still have a few follow-up questions and would appreciate your input.
- In the pick_real task, the TrajGen module uses the prompt “a person stands in place, grabs the cup from side and lifts up”, which is consistent with pick_ub_real. In contrast, pick_sim uses “a person walks to cup, grabs the cup from side and lifts up”.
I noticed that a similar design choice also appears in bimanual_pick. Was this difference intentional? More specifically, does the real setting deliberately exclude the locomotion component, assuming the robot is already positioned near the object?
2.When using ./isaaclab.sh -p scripts/reinforcement_learning/rsl_rl/play_eval.py --task=Isaac-Motion-Tracking--v0, the code instantiates the gym-registered environment Isaac-Motion-Tracking--Real-v0 instead of the expected Isaac-Motion-Tracking--Real-EVAL-v0. This is confusing when the evaluation (inference) environment is intended to differ from the training environment (e.g., whether the REAL environment includes command inputs), as it is not clear which variant is actually used during evaluation. Additionally, there appears to be a minor naming inconsistency: the gym registration uses EVALENV, while the corresponding configuration class is named ENVEVALCFG, which seems to be a small typo.
3.I am still unable to obtain the expected performance when training in the pick_real environment, and the learned behavior shows a significant deviation from the intended motion. I am not sure whether this is due to configuration choices, reward design, observation formulation, or other factors specific to the real setup. Could you please provide some guidance when training pick_real?
4.As a side note, in pick_real, the object-relative term is defined as rel_pos_object = ObsTerm(func=rel_pose_object, params={"fix_height": True, "only_pos": True}), whereas in pick_ub_real, it is defined as rel_pose_object = ObsTerm(func=rel_pose_object, params={"fix_rel_base_height": True, "base_height": 0.753}). I believe they should be unified to a consistent 3-dimensional representation. These appear to be minor typos.
Thank you very much for your detailed response. I really appreciate your time and guidance, and I look forward to further discussions and advice from you.
Best,
ssb
- In the pick_real task, the TrajGen module uses the prompt “a person stands in place, grabs the cup from side and lifts up”, which is consistent with pick_ub_real. In contrast, pick_sim uses “a person walks to cup, grabs the cup from side and lifts up”.
- Yes, correct: The tasks deployed on real g1 in this paper do not involve locomanipulation. The primary reason for this choice is that OWLv2 is used in this paper to obtain object/goal pose observations, which run very slowly (1 Hz). Hence, it became infeasible to run in real-time for tasks that involve locomotion, like walking to pick an object, as the relative object pose will change. One way to tackle this would be to use a real-time object tracker, or SLAM, for localization, which could be used in this specific case. The best way would be to distill the privileged policy to a vision-based closed-loop policy, which is left as follow-up work.
- Yes, you can run play_eval.py with --task=Isaac-Motion-Tracking-Pick-Real-Eval-v0 to run the Pick-real task with sampled object spawn locations instead of instantiating as per the motions used in training. However, because the object locations in training motions are already dense, EnvEvalCfg should yield the same performance as just EnvCfg. Use EnvCfg first to check if training is successful and able to lift most objects instantiated as per training motions, and EnvEvalCfg to check if the trained model also works with sampled locations of objects. There should be very little difference in performance. The names have been corrected. Sorry for the typo
- Try now; there was a small bug while refactoring the code. Should work now, please run training for 4000 iterations
- Yes, correct. They were added to match the training code with the tested checkpoints that can be used directly for inference. Having a unified, consistent representation would be better. Please feel free to unify representations for your applications if you plan to use this code. You are also welcome to open pull requests with cleaner code and new, tested checkpoints
Best,
Dvij- added a commit that references this issue
on Jul 16, 2026
Metadata
Metadata
Assignees
Labels
No labels
Thank you for open-sourcing such an amazing project! While attempting to reproduce the code, I encountered a few issues and would appreciate your clarification:
During the training of pick-real, the environment does not initialize a table, causing the target object to fall directly to the ground. This makes it impossible for the robot to grasp the object. Is this a bug or an intentional training feature?
In pick-real, the observation is 100-dimensional, with rel_pose_object comprising 7 dimensions (3 for position and 4 for quaternion orientation). However, in the MuJoCo environment, the observation is 96-dimensional, and rel_pose_object only includes 3 dimensions (position, without orientation). This discrepancy prevents the weights trained in pick-real from running in MuJoCo. Is this intentional, or is there an updated version of pick-real? Additionally, the deploy_pick script in MuJoCo also encounters the issue of objects falling, making grasping impossible.
In motion_lib, when loading reference motion data, a motion_offset is used to define a positional bias. Specifically, the x-axis of the grasping position is set to 2.1. Does this mean the robot can only grasp objects at x = 2.1, or is there another reason for this value? Given that pick-real currently does not train correctly, this is quite confusing.
The paper and code do not clearly explain how the robot generalizes to grasping objects at positions different from the reference motion. Could you provide some insights or demos on this?
The project includes HumanoidVerse, and the demo interface also shows some demos from HumanoidVerse. However, the README file does not cover how to reproduce or use these parts. Could you please make these methods publicly available?
Thanks for sharing this great project. Merry Christmas! Bese wishes.