Update Experiment Group Metric Settings
Set per-group ranking weights. Replaces the full configuration; None = inherit project-level defaults.
Path parameters
Request body
Response
Successful Response
Changes
Changed in 3 of the 14 revisions of this API.51622
- ▲
removed the enum value
Measures which documents or chunks retrieved were used by the model to generate a responseand how much of the text in the retrieved chunks was used by the model to compose its response.of the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-removed
- ○
removed the
Measures which documents or chunks retrieved were used by the model to generate a responseand how much of the text in the retrieved chunks was used by the model to compose its response.enum value from themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-removed
This revision also has 44 changes that name no endpoint, such as unreferenced schemas being removed. See the revision's changelog
- ▲
- ●
added the new
Assesses whether a chatbot interaction that includes audio left the user feeling satisfied and positiveor frustrated and dissatisfiedbased on toneengagementand overall experience.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Assesses whether a chatbot interaction that includes images or documents left the user feeling satisfied and positiveor frustrated and dissatisfiedbased on toneengagementand overall experience.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures how sexist audio input content might be perceivedranging from 0 to 1.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures how sexist image or document input content might be perceivedranging from 0 to 1.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures how well the LLM follows system instructions provided in the prompt.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures how well the LLM follows system instructions provided in the prompt.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ○
added the new
Assesses whether a chatbot interaction that includes audio left the user feeling satisfied and positiveor frustrated and dissatisfiedbased on toneengagementand overall experience.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Assesses whether a chatbot interaction that includes images or documents left the user feeling satisfied and positiveor frustrated and dissatisfiedbased on toneengagementand overall experience.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures how sexist audio input content might be perceivedranging from 0 to 1.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures how sexist image or document input content might be perceivedranging from 0 to 1.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures how well the LLM follows system instructions provided in the prompt.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures how well the LLM follows system instructions provided in the prompt.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
This revision also has 8 changes that name no endpoint, such as unreferenced schemas being removed. See the revision's changelog
- ●
- ▲
removed the enum value
A measure of the model's own confusion in its output. Higher scores indicate higher uncertainty.of the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-removed
- ▲
removed the enum value
BLEU is a case-sensitive measurement of the difference between an model generation and target generation at the sentence-level.of the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-removed
- ▲
removed the enum value
Measures the perplexity of the prompt. Lower perplexity score is generally considered to be better because it means the model is less surprised by the text and can predict the next word in a sentence with higher accuracy.of the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-removed
- ▲
removed the enum value
ROUGE measures the unigram overlap between model generation and target generation as a single F-1 score.of the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-removed
- ●
added the new
Detects a significant shift in the user's primary conversational goal or workflow during a session that includes audio.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Detects a significant shift in the user's primary conversational goal or workflow during a session that includes images or documents.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Detects whether the user successfully accomplished all of their goals in a session that includes audio.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Detects whether the user successfully accomplished all of their goals in a session that includes images or documents.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures how well the workflow's response aligns with ground truth that includes audio.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures how well the workflow's response aligns with ground truth that includes images or documents.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures the coherence of the reasoning process by evaluating the consistency and logical flow of reasoning steps with audio context.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures the coherence of the reasoning process by evaluating the consistency and logical flow of reasoning steps with image or document context.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures the potential presence of factual errors or inconsistencies in the model's responseincluding audio content.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ●
added the new
Measures the potential presence of factual errors or inconsistencies in the model's responseincluding image or document content.enum value to themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-added
- ○
the endpoint scheme security
ClassicAPIKeyHeaderwas added to the APIapi-security-added
- ○
added the new
Detects a significant shift in the user's primary conversational goal or workflow during a session that includes audio.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Detects a significant shift in the user's primary conversational goal or workflow during a session that includes images or documents.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Detects whether the user successfully accomplished all of their goals in a session that includes audio.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Detects whether the user successfully accomplished all of their goals in a session that includes images or documents.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures how well the workflow's response aligns with ground truth that includes audio.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures how well the workflow's response aligns with ground truth that includes images or documents.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures the coherence of the reasoning process by evaluating the consistency and logical flow of reasoning steps with audio context.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures the coherence of the reasoning process by evaluating the consistency and logical flow of reasoning steps with image or document context.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures the potential presence of factual errors or inconsistencies in the model's responseincluding audio content.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
added the new
Measures the potential presence of factual errors or inconsistencies in the model's responseincluding image or document content.enum value to the request propertymetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/request-property-enum-value-added
- ○
removed the
A measure of the model's own confusion in its output. Higher scores indicate higher uncertainty.enum value from themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-removed
- ○
removed the
BLEU is a case-sensitive measurement of the difference between an model generation and target generation at the sentence-level.enum value from themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-removed
- ○
removed the
Measures the perplexity of the prompt. Lower perplexity score is generally considered to be better because it means the model is less surprised by the text and can predict the next word in a sentence with higher accuracy.enum value from themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-removed
- ○
removed the
ROUGE measures the unigram overlap between model generation and target generation as a single F-1 score.enum value from themetric_weights_configuration/anyOf[subschema #1]/additionalProperties/description/anyOf[subschema #1: MetricDescriptions]/response property for the response status200response-property-enum-value-removed
This revision also has 15 changes that name no endpoint, such as unreferenced schemas being removed. See the revision's changelog
- ▲