---
title: "Reward model score mathematical notation"
description: "Reward model score is a recurring research-paper notation family. Maps a prompt-response pair or trajectory to a learned scalar used as a proxy for human…"
canonical_url: "https://fanout.sh/labs/math-decoder/symbol/reward-model"
md_url: "https://fanout.sh/labs/math-decoder/symbol/reward-model.md"
last_updated: "2026-08-09"
access: "public"
---

# Reward model score mathematical notation

Reward model score is a recurring research-paper notation family. Maps a prompt-response pair or trajectory to a learned scalar used as a proxy for human…

## Public overview

Reward model score is a recurring research-paper notation family. Maps a prompt-response pair or trajectory to a learned scalar used as a proxy for human preference.

Reward model score: Maps a prompt-response pair or trajectory to a learned scalar used as a proxy for human preference. Example: Score the preferred response with the learned reward network.

---
This representation contains public Fanout content only. Protected Pro lessons, account data, billing, checkout, and pricing are not included.

Browse the public content map: https://fanout.sh/sitemap.md
