VLA fine-tuning method

OpenVLA-OFT

OpenVLA-OFT is an optimization and fine-tuning recipe that changes action decoding and representation for OpenVLA, with a separately documented bimanual extension.

Source contextPrimary specReviewed
Sources 3
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

OpenVLA-OFT research project

Moo Jin Kim, Chelsea Finn and Percy Liang · Publication date not disclosed · Source accessed 2026-09-28

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Reviewed recordLast reviewed 3 sources ↗4 sourced fields

Architecture and scope

Parallel decoding, action chunking and a continuous L1 objective form the core recipe. The OFT+ extension adds feature modulation for the ALOHA setting; it is not the identical configuration behind every LIBERO row.

Source contextPrimary specReviewed
Sources 3
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

OpenVLA-OFT research project

Moo Jin Kim, Chelsea Finn and Percy Liang · Publication date not disclosed · Source accessed 2026-09-28

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Integration and evaluation

The imported LIBERO suite results use the final Table I configuration with wrist images and proprioception. Each suite is fine-tuned independently, so the results are not zero-shot generalization.

Source contextPrimary specReviewed
Sources 3
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

OpenVLA-OFT research project

Moo Jin Kim, Chelsea Finn and Percy Liang · Publication date not disclosed · Source accessed 2026-09-28

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Evidence boundary

Paper results are author-reported and were not reproduced by PhysicalAI.best.

Source contextPrimary specReviewed
Sources 3
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

OpenVLA-OFT research project

Moo Jin Kim, Chelsea Finn and Percy Liang · Publication date not disclosed · Source accessed 2026-09-28

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Model identity

Developer
Moo Jin Kim, Chelsea Finn and Percy Liang
Primary specReviewed
Sources 1
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Exact version / scope
OpenVLA-OFT; arXiv:2502.19645v1 result profile
Primary specReviewed
Sources 1
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Category
VLA fine-tuning method
Primary specReviewed
Sources 1
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Access model
Public code and task-specific checkpoints
Primary specReviewed
Sources 1
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

License
Not publicly disclosed

Inputs, outputs & embodiment

Inputs
Not publicly disclosed
Outputs
Not publicly disclosed
Embodiment
Not publicly disclosed

Implementation & training

Compute
Not publicly disclosed
Training-data disclosures
Not publicly disclosed
Integrations
Not publicly disclosed

Reported benchmark results

Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
Model / revisionBenchmark / metricReported resultEnvironmentEvidence & scope
OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Spatial97.6Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Object98.4Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Goal97.9Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Long94.5Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

What this evidence does not establish

  • Paper results are author-reported and were not reproduced by PhysicalAI.best.
  • Pricing and deployment service terms were not verified.

Relationships & deployments

Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.

evaluated on
OpenVLA-OFT → LIBERO
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Connection reviewed 2026-09-28

uses
OpenVLA-OFT → OpenVLA
Sources 1
moojink/openvla-oft official repository

moojink · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Connection reviewed 2026-09-28

Record history & verified changes

A research review records when we checked a source. It does not mark a product launch or a new deployment.

Research review / additionRegistry addition: exact paper revision, task splits, simulated embodiment and reuse terms recorded.

Four contextualized LIBERO suite results added

Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Inspect the evidence

Sources & evidence

md-oft-project
OpenVLA-OFT research project

Moo Jin Kim, Chelsea Finn and Percy Liang · Paper · Publication date not disclosed · Source accessed 2026-09-28

Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.