Inference AIops
本地Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.
安装配置
{
"mcpServers": {
"inference-aiops": {
"command": "npx",
"args": [
"inference-aiops"
]
}
}
}
Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.
{
"mcpServers": {
"inference-aiops": {
"command": "npx",
"args": [
"inference-aiops"
]
}
}
}