Plugin Name: EvalPort Result Export
Description:
Hi Portkey team — I maintain EvalPort (https://github.com/adhabnr-ux/evalport), an open JSON-Schema spec for portable LLM eval test cases, graders, and result sets, with Python/TS SDKs that validate against the schemas.
Looking at plugins/types.ts and the Patronus checks (e.g. plugins/patronus/isHelpful.ts), every guardrail plugin already returns a PluginHandlerResponse with a verdict: boolean and a data payload — which is essentially a grader result waiting to be serialized. isHelpful.ts even calls its check evaluator = 'judge'.
Idea for a small plugin (or gateway-level hook) that collects the PluginHandlerResponse[] produced across a request's beforeRequestHook/afterRequestHook checks and emits them as an EvalPort ResultSet (schema: https://github.com/adhabnr-ux/evalport/blob/main/spec/schemas/resultset.json), e.g.:
{
"test_case_id": "<request id>",
"grader_results": [
{ "grader_id": "patronus:is-helpful", "type": "llm-judge", "score": 1, "passed": true },
{ "grader_id": "patronus:no-gender-bias", "type": "llm-judge", "score": 1, "passed": true }
],
"passed": true
}
The concrete use case: teams running EvalPort-defined TestCase/EvalSuite fixtures (https://github.com/adhabnr-ux/evalport/blob/main/spec/schemas/testcase.json) through the gateway's own guardrail configs offline, to regression-test a guardrail setup the same way they'd regression-test a model — and get back a portable, diffable result artifact rather than one tied to a single guardrail provider's dashboard.
Full spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
Following the plugin proposal format from plugins/Contributing.md — happy to build this out as an actual plugin under /plugins/evalport and open a PR if there's interest. No pressure if this doesn't fit the roadmap.
Plugin Name: EvalPort Result Export
Description:
Hi Portkey team — I maintain EvalPort (https://github.com/adhabnr-ux/evalport), an open JSON-Schema spec for portable LLM eval test cases, graders, and result sets, with Python/TS SDKs that validate against the schemas.
Looking at
plugins/types.tsand the Patronus checks (e.g.plugins/patronus/isHelpful.ts), every guardrail plugin already returns aPluginHandlerResponsewith averdict: booleanand adatapayload — which is essentially a grader result waiting to be serialized.isHelpful.tseven calls its checkevaluator = 'judge'.Idea for a small plugin (or gateway-level hook) that collects the
PluginHandlerResponse[]produced across a request'sbeforeRequestHook/afterRequestHookchecks and emits them as an EvalPortResultSet(schema: https://github.com/adhabnr-ux/evalport/blob/main/spec/schemas/resultset.json), e.g.:{ "test_case_id": "<request id>", "grader_results": [ { "grader_id": "patronus:is-helpful", "type": "llm-judge", "score": 1, "passed": true }, { "grader_id": "patronus:no-gender-bias", "type": "llm-judge", "score": 1, "passed": true } ], "passed": true }The concrete use case: teams running EvalPort-defined
TestCase/EvalSuitefixtures (https://github.com/adhabnr-ux/evalport/blob/main/spec/schemas/testcase.json) through the gateway's own guardrail configs offline, to regression-test a guardrail setup the same way they'd regression-test a model — and get back a portable, diffable result artifact rather than one tied to a single guardrail provider's dashboard.Full spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
Following the plugin proposal format from
plugins/Contributing.md— happy to build this out as an actual plugin under/plugins/evalportand open a PR if there's interest. No pressure if this doesn't fit the roadmap.