Async inference is the part I keep pointing people to. Has anyone on the team looked at how SmolVLA follows instructions phrased unlike the training data? In LIBERO I saw ordinary paraphrases do almost nothing while an odd framing did a lot, and I'd like to compare notes with whoever owns evals.
Sattyam Jain
Sattyam
·
AI & ML interests
Multi-Agent Systems, AgentOps, LLM Security, Agent Memory, MCP, Production AI Infrastructure
Recent Activity
new activity 3 days ago
openEuler/pi05:Any eval for this port? new activity 3 days ago
hi-space/PI-0.5-Pick-Banana-v2:What changed from v1 to v2? new activity 3 days ago
Cache-SCA/pi05_teleop_close_pot:When does close-pot count as a success?Organizations
Any eval for this port?
#7 opened 3 days ago
by
Sattyam
What changed from v1 to v2?
#1 opened 3 days ago
by
Sattyam
When does close-pot count as a success?
#1 opened 3 days ago
by
Sattyam
What does eval1 hold out?
#1 opened 3 days ago
by
Sattyam
Real Franka or sim?
#1 opened 3 days ago
by
Sattyam
Five-task model on the cup task?
#1 opened 3 days ago
by
Sattyam
How was the tea task scored?
#1 opened 3 days ago
by
Sattyam
commented on SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data 3 days ago
Does colour wording matter for sort-block?
#1 opened 3 days ago
by
Sattyam
Task and eval for this fine-tune?
#2 opened 5 days ago
by
Sattyam
Eval instructions for the RoboTwin checkpoint
#2 opened 5 days ago
by
Sattyam
One trial per perturbed task: is there an interval on each column?
#2 opened 6 days ago
by
Sattyam
Good to see on-device numbers for VLA fine-tunes. After the on-device optimizations, did you compare the optimized policy with the original on the same episodes, including the failed ones? A port can keep its success rate and still move somewhere new when it fails, and on an embedded robot that's the part that matters.
commented on Generalist Robot Policy Evaluation in Simulation with NVIDIA Isaac Lab-Arena and LeRobot 6 days ago
EnvHub for evaluation makes a lot of sense. Would it also be a good place for shared perturbation sets, like a fixed list of reworded instructions per task with the original as the control? Right now everyone writes their own, so wording results from different groups can't be compared.
Any eval runs on this one?
#1 opened 6 days ago
by
Sattyam
pi0.5 vs SmolVLA on the same task?
#1 opened 6 days ago
by
Sattyam
How did you pick this step?
#1 opened 6 days ago
by
Sattyam