Robotics Institute · Carnegie Mellon University

Narayanan
Palghat Parameswaran

I work on robot learning.

From the lab01 / 08
Unitree G1 · simulationWhole-body box carry, generated from a text prompt
About

Hi, I'm Narayanan.

Portrait of Narayanan Palghat Parameswaran
Pittsburgh, PA

How can robots pick up new physical skills from people, from language, and from their own experience? That question runs through most of my work, from generative models for control to whole-body humanoid skills.

I'm a Master's student in Robotic Systems Development at Carnegie Mellon. With Deva Ramanan, I study how language models can teach robots new skills: you describe a motion, and the robot learns to perform it. At Sebastian Scherer's AirLab, I'm bringing the same idea to humanoids, learning whole-body skills from human demonstrations.

This past summer at FieldAI, I worked on curricula that let quadrupeds learn parkour from their own failures. Before CMU, I spent two years with Shishir Kolathaya at IISc on diffusion policies for multi-skill locomotion, and along the way collaborated with Pulkit Agrawal (MIT) on sim-to-real.

Generative control policiesWhole-body controlTest-time steeringDiffusion policiesVision-language-action models
News

Recent updates

Papers, moves and milestones since 2024.

  • Paper

    FCTS will appear at the IROS 2026 SARL workshop. The project page is live.

  • Preprint

    MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation is on arXiv, with a project page and hardware videos.

  • Lab

    Joined AirLab at CMU as a Graduate Student Researcher with Sebastian Scherer, working on humanoid loco-manipulation.

  • Internship

    Wrapped up my research internship at FieldAI.

  • Internship

    Started as a Robotics Research Intern at FieldAI in Irvine, CA, on perceptive quadruped parkour.

  • Award

    Third place in CMU 10-623 (Generative AI) for inference-time diffusion steering with a GRPO-style reward model.

  • Paper

    MimicAgent presented at the ICLR 2026 RSI workshop (OpenReview).

  • CMU

    Started the MS in Robotic Systems Development at Carnegie Mellon's School of Computer Science and joined Deva Ramanan's lab.

  • Admits

    Admitted to MS programs at CMU, Georgia Tech, UPenn and UMD. Chose CMU.

  • Preprint

    GRoQ-LoCO, a generalist and robot-agnostic quadruped locomotion policy learned from offline data, is on arXiv.

  • Paper

    HySem presented at the NeurIPS 2024 Table Representation Learning workshop.

  • Lab

    Became a full-time Research Assistant at the Stochastic Robotics Lab, IISc Bangalore.

  • Degree

    Graduated with a B.Tech in Mechanical Engineering from NIT Bhopal.

  • Lab

    Joined the Stochastic Robotics Lab, IISc Bangalore as a Research Intern with Shishir Kolathaya.

Research

Selected publications

* denotes equal contribution.

FieldAI × CMU FCTS policy reaching all waypoints on a scanned debris site where Extreme Parkour and Eurekaverse baselines fail
IROS 2026 / SARL Workshop

FCTS: Failure-Guided Curriculum Tree Search Over Generated Terrain for Legged Parkour

Narayanan P. Parameswaran, D. K. Kim, J. Patrikar, …, S. Scherer

When the policy fails on a course, FCTS records where and how it failed, and an LLM uses that to generate new terrain aimed at exactly that weakness. The policy then trains on it, and Monte Carlo Tree Search decides which of several competing curricula deserves more of the compute budget.

85.8%success on medium routes, up from 28.4%
44.7%success on hard routes, up from 14.9%
5.3×faster than sequential search
Deva's Lab · CMU
ICLR 2026 / RSI Workshop · arXiv 2026

MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation

Lucky Kant Nayak*, Narayanan P. Parameswaran*, Neehar Peri, Deva Ramanan (*equal contribution)

Writing a plausible reference motion is easier than writing a robust reward. Coding agents turn a prompt such as "wheeled quadruped doing a frontflip" into a kinematic trajectory, check it with task-agnostic unit tests, and repair the code until it passes. DeepMimic-style RL then makes it physical.

87%of prompts yield semantically aligned references
92%of trained policies match their prompt
43×fewer LLM tokens than Eureka
@article{nayak2026mimicagent,
  title   = {MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation},
  author  = {Nayak, Lucky Kant and Parameswaran, Narayanan Palghat and Peri, Neehar and Ramanan, Deva},
  journal = {arXiv preprint arXiv:2609.24145},
  year    = {2026}
}
IISc Bangalore
arXiv 2025

GRoQ-LoCO: Generalist and Robot-agnostic Quadruped Locomotion Control using Offline Datasets

Narayanan PP, Sarvesh Prasanth Venkatesan, Srinivas Kantha Reddy, Shishir Kolathaya

One policy, many quadrupeds. A recurrent transformer is trained by offline behavior cloning on 2,000+ expert trajectories per morphology. It fuses stair and flat-ground skills from proprioception alone, with no robot-specific encodings, and runs on a Unitree Go1 and the 20 kg-payload Stoch 5 without fine-tuning.

5+unseen morphologies, zero-shot
0robot-specific encodings or fine-tuning
@article{pp2025groqloco,
  title   = {GRoQ-LoCO: Generalist and Robot-agnostic Quadruped Locomotion Control using Offline Datasets},
  author  = {PP, Narayanan and Venkatesan, Sarvesh Prasanth and Reddy, Srinivas Kantha and Kolathaya, Shishir},
  journal = {arXiv preprint arXiv:2505.10973},
  year    = {2025}
}
JN Research Labs HySem pipeline diagram
NeurIPS 2024 / TRL Workshop

HySem: A Context Length Optimized LLM Pipeline for Unstructured Tabular Extraction

Narayanan PP, Anantharaman Palacode Narayana Iyer

Tokenizer-aligned context compression lets an on-premise 8B LLM turn messy pharmaceutical HTML tables, with nested headers and domain terms, into semantic JSON.

−39%context length, lossless
+25 ptsover open-source baselines
@inproceedings{pp2024hysem,
  title     = {HySem: A context length optimized LLM pipeline for unstructured tabular extraction},
  author    = {PP, Narayanan and Iyer, Anantharaman Palacode Narayana},
  booktitle = {NeurIPS 2024 Third Table Representation Learning Workshop},
  year      = {2024}
}

Technical reports

2024MorphDLoco: Morphology-aware Diffusion Policy for Multi-robot Multi-skill LocomotionTechnical report · IIScPDF
2024Diverse Quadruped Locomotive Behavior on Different TerrainsB.Tech project report · NIT BhopalPDF
2023End-to-end Reinforcement-based Quadruped Locomotion in PyBulletInternship report · IIScPDF
2022Designing Joystick Control Interfaces for Service Robot TeleoperationInternship report · JN Research LabsPDF
Robots in action

Quadrupeds to humanoids

Legged robots on hardware, and humanoid skills in simulation.

Experience

Where I've worked

  1. Sep 2026 — now
    ● current

    AirLab, Carnegie Mellon University

    Graduate Student Researcher · Advisor: Sebastian Scherer · Pittsburgh, PA

    Learning whole-body humanoid skills from human demonstrations.

  2. May — Aug 2026

    FieldAI

    Robotics Research Intern · Irvine, CA

    Automatic curricula for perceptive quadruped parkour.

  3. Aug 2025 — now
    ● current

    Deva's Lab, Carnegie Mellon University

    Graduate Student Researcher · Advisor: Deva Ramanan · Pittsburgh, PA

    Teaching legged robots new skills from language.

  4. May 2023 — Jun 2025

    Stochastic Robotics Lab, IISc Bangalore

    Research Intern (May 2023 – Jun 2024) · Research Assistant (Jul 2024 – Jun 2025) · Advisor: Shishir Kolathaya

    Multi-skill diffusion policies and generalist quadruped locomotion, with a sim-to-real collaboration with Pulkit Agrawal (MIT).

  5. May 2022 — Jul 2024

    JN Research Labs

    Research Intern · Summer Intern · Mentor: Anantharaman P. N. · Bangalore

    LLMs for document understanding.

Projects

Things I've built

🏆 3rd place
CMU 10-6232026

Inference-Time Diffusion Steering with a GRPO-Style Reward Model

A timestep-conditioned DiT reward model, trained with a Bradley–Terry objective on group rollouts, steers a frozen Motion Diffusion Model at every denoising step. Physics violations drop 11%.

DiffusionReward modelsMDM
View project
CMU 16-6622026

Open-Vocabulary Robotic Sorting with Vision-Language Models

A zero-shot pipeline with a VLM in the loop for a Franka Panda. It grounds open-vocabulary sorting rules in a structured memory of the scene state and reaches 90% placement accuracy across 3 VLMs.

Franka PandaVLMsManipulation
View project
Research CoPilot agent graph
Open source2025

Research CoPilot: Multi-Agent Auto-Research

Planner, retriever, critic and writer agents with RAG over retrieved literature. They turn raw ideas into grounded, publication-ready research proposals.

LangGraphRAGMulti-agent
GitHub
Multi-skill diffusion policy diagram
IISc2025

Multi-Skill Diffusion Policies for Generalist Quadruped Control

A conditional DDPM policy that fuses diverse gaits into one generalist controller. It generates multi-skill actions in real time and transfers zero-shot to 5+ unseen morphologies and terrains.

DDPMBehavior cloningCross-embodiment
View project
IISc · B.Tech thesis2024

QuadScape: Diverse Gaits on Diverse Terrains

Asymmetric actor-critic RL with terrain-dependent auxiliary rewards. Trotting, pronking and bounding over stairs and slopes in Isaac Gym, deployed on a Unitree Go1.

Isaac GymPPOGo1
Read report
IISc2023

StochBullet: End-to-End RL Locomotion in PyBullet

A reproduction of RMA's teacher–student training with PPO and a fixed curriculum on the Stoch 3 quadruped. It stays robust to pushes and friction changes.

PyBulletRMAPPO
Read report
DepthForge
NIT Bhopal2023

DepthForge: Monocular Distance as Classification

A Webots data pipeline with a camera and stacked-LiDAR rover, plus a temporal CNN over 5-frame sequences. Over 90% accuracy in an unseen indoor scene, and it transfers to the real world at close range.

WebotsTemporal CNN
Architecture
CaptionCraft
Independent2024

CaptionCraft: Region-Aware Image Captioning

DETR boxes and MiniGPT-4 region captions, fused by Mistral-7B into denser, better-grounded image descriptions.

DETRMiniGPT-4Mistral
GitHub
Four-wheeled differential-drive indoor robot
JN Research Labs2022

Indoor Obstacle Avoidance with Actor-Critic RL

Actor-critic RL policies for obstacle avoidance on a four-wheeled differential-drive robot in indoor environments. The accompanying Indoor Robot Navigation Dataset won a Kaggle bronze medal.

Actor-critic RLNavigationDataset
Dataset
Background

Education & honors

Education

Carnegie Mellon University

2025 — May 2027

M.S. Robotic Systems Development, School of Computer Science

GPA 4.07 / 4.0

National Institute of Technology, Bhopal

2020 — 2024

B.Tech. Mechanical Engineering

GPA 3.75 / 4.0

Honors & awards

  • Third Place, CMU 10-623 Generative AI course project
  • MS admits: CMU, Georgia Tech, UPenn, UMD
  • NPTEL Topper in three courses offered by the IITs
  • Kaggle Bronze Medal, Indoor Robot Navigation Dataset
  • 5th place, Vision Verse CV hackathon, IIT Mandi
  • Certificate of Appreciation, Shaastra '22, IIT Madras

Programming & ML

PythonPyTorchHuggingFaceLangChainLangGraphMATLAB

Distributed training

PyTorch DDPDeepSpeedRayAMPtorch.compile

Robotics & simulation

Isaac LabIsaac GymMuJoCo PlaygroundPyBulletWebotsGazebo

Tools & DevOps

KubernetesLinuxAWSDockerGit