Purpose
Operate archive search from the caller's host. Compose and Kubernetes use the
same vss configure and vss search run commands; only the deployment origin
differs. Source ingestion and deletion are Agent-backed when the deployment has
an agent /api route; on a build without one, they belong to
vss-manage-video-io-storage references/provision-vios-source.md.
Hard boundaries
- Run the project-local CLI on the host. Never use
docker exec,kubectl exec, a pod shell, or a globally installedvssas a substitute. - Never improvise a mutation against Elasticsearch, RTVI-CV, RTVI-Embed,
storage-ms, or VST. Two paths are sanctioned, and the deployment picks which:
the Agent upload/delete lifecycle where an agent
/apiroute answers, andvss-manage-video-io-storagereferences/provision-vios-source.mdwhere none does. That recipe owns the direct calls this rule otherwise forbids. - Never remove, broaden, or silently substitute a requested source constraint.
- Similarity is retrieval evidence, not proof of visual presence.
- The CLI attempts critic verification by default. Do not separately inspect screenshots or call another verifier during the initial search turn.
- Offer delegated verification only when every displayed result is
unverified, and only after displaying them and receiving explicit user confirmation. If any result isconfirmedorrejected, do not hand off any result to another verifier.
Prerequisites
- A running VSS
searchprofile and its host-reachable Compose or Ingress origin. - A checkout containing
libs/vss, hostuv,curl, andjq. vss vios listfor source listing and inspection (same CLI, same recorded origin).
Resolve and validate the checkout once:
bashVSS_REPO_ROOT="${VSS_REPO_ROOT:-$HOME/video-search-and-summarization}" test -f "${VSS_REPO_ROOT}/libs/vss/pyproject.toml" || { echo "VSS checkout not found at ${VSS_REPO_ROOT}; set VSS_REPO_ROOT explicitly" >&2 exit 1 } VSS=(uv run --project "${VSS_REPO_ROOT}/libs/vss" vss) cd "${VSS_REPO_ROOT}" && "${VSS[@]}" search run --help >/dev/null || exit 1
libs/vss is the library's own workspace, so no extras and no --no-dev are
needed — the agent stack is not in it.
Resolve the deployment through its one public/host origin:
bashif [ -z "${VSS_ORIGIN:-}" ]; then VSS_ORIGIN=$("${VSS[@]}" configure show 2>/dev/null | jq -er '.base_url | select(type == "string" and length > 0)') || { echo "Provide the Compose or Ingress origin" >&2 exit 1 } fi VSS_ORIGIN="${VSS_ORIGIN%/}" VST_URL="${VSS_ORIGIN}" VSS_VIOS_URL="${VSS_ORIGIN}/vst" "${VSS[@]}" configure --base-url "${VSS_ORIGIN}" || exit 1
In a persisted multi-step workflow, reuse the origin recorded by the prepared deployment as above. Do not repeat public-origin selection, edit routing, or redeploy merely because the next agent turn did not inherit shell variables.
See deployment resolution
for the deployment-owned VSS_PUBLIC_URL contract. On Kubernetes, never use
port-forwarding, Service DNS, NodePorts, or a guessed Helm release. Routes not
exposed through the Ingress are recorded as absent and a search path needing
one exits 4.
For deployment readiness, ingestion, fixture cleanup, index checks, RTSP, or
deletion, read source lifecycle completely
before acting. Re-run vss configure after the first ingestion: the recorded
raw family is what enables frame-level lookups (it gates frames_index, which
attribute and fusion need for frame enrichment). Only source-type selection is
independent of the index inventory.
Mandatory search workflow
-
Confirm the selected deployment is the
searchprofile. If required routes are unavailable, ask whether to reconnect or deploy it withthe/vss-build-vision-aistock Search workflow; do not target another profile. -
When the user names a file, camera, or sensor, list registered sources with
"${VSS[@]}" vios listbefore invoking the search CLI — it reads the originvss configurerecorded, so it takes no endpoint. Accept only an exact source, stream ID, or one unambiguous normalized substring match.- No match: report the missing source, list available names, and ask the
user to clarify or explicitly request ingestion. Stop without probing the
search CLI, deploying, or ingesting. Never continue with a different
source. Answering about
warehouse_samplewhen the request namedwarehouse-ladderreturns a confident answer about the wrong video, and nothing downstream can tell it was substituted. - Several matches: ask the user to choose and stop.
- Never substitute another video or run an unrestricted search as a probe.
Preserve both the matched source's
.sensorIdand.name. The--video-sourcevalue depends on the search path, not the source type (optional for every path):embedmatches the sensor ID literally;attributeandobjectmatch the name literally; onlytagresolves a source name to its VST sensor ID (passing an already-id through).fusiondoes not resolve — its embedding leg filters by sensor ID literally — so hand fusion the preserved sensor ID (the tag leg accepts IDs too). For every path an unknown source yields an empty, narrowed result, not an error. Set--source-type video_filefor uploads or--source-type rtspfor live streams. This selects the index partition for that media kind from a fixed uploads anchor (not a discovered index), independently of the identifier, so it is correct regardless of ingestion order. - No match: report the missing source, list available names, and ask the
user to clarify or explicitly request ingestion. Stop without probing the
search CLI, deploying, or ingesting. Never continue with a different
source. Answering about
-
Decompose the request before choosing a path; do not pick by surface form.
run embedaccepts any sentence, so being one sentence is not evidence for embed. Separate each specific detectable property (white jacket,red hard hat) from the actions/relations only embeddings capture, then choose:- a detectable property plus an action or relation is present →
run fusion(even within one sentence) - free-text intent with no detectable property →
run embed - detectable properties only, no action or relation →
run attribute - explicit tracked object IDs →
run object - explicit keyword or tag intent — lexical (BM25) match against indexed VLM tags, with no detectable property and no semantic free-text →
run tag
--attributeis for specific detectable properties, not generic nouns or actions. A property counts only when RT-CV detects it on the subject (attire, PPE, color-on-person), not object identity or an object's own color; keepred forkliftwholly in--query.worker in a hard hat carrying a conehas a property (hard hat) and an action (carrying a cone):run fusion --query "worker in a hard hat carrying a cone" --attribute "hard hat". Reserve embed for genuinely attribute-free intent.run tagis for explicit lexical intent — matching indexed VLM tag keywords by BM25 — not semantic similarity; reserve it for keyword/tag queries that name no detectable property. - a detectable property plus an action or relation is present →
-
Construct the invocation as a Bash array and validate only its exact stdout. Read CLI usage for every supported flag.
bash: "${SEARCH_PATH:?set embed|attribute|fusion|object|tag}" : "${SOURCE_TYPE:?set video_file or rtsp}" TOP_K="${TOP_K:-3}" VIDEO_SOURCES=() # sensor IDs for embed/fusion; names for attribute/object/tag : "${SOURCE_SCOPED:?set true for a resolved scope; false only when unrestricted}" if [ "${SOURCE_SCOPED}" = true ] && [ "${#VIDEO_SOURCES[@]}" -eq 0 ]; then echo "Resolved source scope is empty; refusing an unrestricted search" >&2 exit 1 fi SEARCH_COMMAND=( "${VSS[@]}" search run "${SEARCH_PATH}" --source-type "${SOURCE_TYPE}" --top-k "${TOP_K}" --raw ) for source in "${VIDEO_SOURCES[@]}"; do SEARCH_COMMAND+=(--video-source "${source}") done # Append --query, repeatable --attribute, --object-id, and time bounds as needed. if ! SEARCH_JSON=$("${SEARCH_COMMAND[@]}"); then echo "Search command failed" >&2 exit 1 fi printf '%s' "${SEARCH_JSON}" | jq -e 'type == "object" and (.data | type == "array")' >/dev/null || { echo "Search did not return a SearchOutput object with a data array" >&2 exit 1 }
Do not pass endpoint, index, model, deployment, profile, or base-URL flags to
search run; vss configure owns those values. Do not replace a failed CLI
call with /api/v1/search or private backend access.
-
Validate each nonempty hit's exact returned
screenshot_urlwith a bounded GET for availability only. Its normalized scheme, host, and effective port always match the origin recorded byvss configure, because the CLI stamps that origin into every hit — a localhost media URL means the deployment was configured against a localhost origin, not that the URL is malformed. On Brev, prefer the public HTTPS secure-link origin. If setup used the documented host-reachable fallback after its one bounded public probe failed, accept only that exact recorded origin and label its media URLs host-local; do not restart routing diagnosis. Reject credentials in the URL and never rewrite the URL or add astreamIdrouting header. Discard the response body; availability is not visual evidence. -
Read every hit's
verificationobject:confirmed: the critic found all requested visual criteria in that clip.rejected: the critic found a visual criterion was not met.unverified: no usable critic verdict was produced. This includes a missing VLM, inaccessible media, and malformed or inconclusive output.
The CLI is fail-open: verification failure must not discard or fail retrieval.
Never derive a verdict from similarity, filenames, object IDs, or screenshot
availability. Treat boolean criteria_met values as critic evidence only.
- Format nonempty results without raw JSON:
text## Video Search Results <each hit's exact source, start/end, similarity, complete media URL, verification result, and criteria when present> Similarity scores are retrieval evidence; the separate verification result records whether the bounded clip satisfied the visual request. ## Verification Step Would you like me to verify the unverified search results?
Include ## Verification Step only when the nonempty displayed result set is
entirely unverified. If any displayed result is confirmed or rejected,
omit it even when other hits are unverified. Never deploy a VLM or call
vss-ask-video automatically during this results turn.
-
If the user explicitly confirms, read search-result verification completely and delegate the displayed hits only after confirming again that every one is still
unverified. Preserve their exact bounded intervals and the complete original visual intent. Keep at most three delegations in flight. Never hand off a partially verified result set. -
If
.datais empty, report zero candidates faithfully — a fact about retrieval, not about the video. Do not claim the object is absent, describe what the footage contains, or argue it is not something you would expect there: a threshold or embedding gap yields the same empty result as a genuine absence. Offer a specific query or similarity-threshold refinement while preserving the source. Never broaden the search silently.
Natural-language Agent responses
Use the host CLI for deterministic structured search. If a caller explicitly
requires the deployment Agent to decompose a natural-language request, its
/api/v1/search response is conversational text, not SearchOutput. Validate
the known text field and present it as prose; never run .data[], screenshot,
or verification parsing against that response or invent structured hit rows.
Troubleshooting
- CLI unavailable: verify
VSS_REPO_ROOTpoints at the checkout, and stop. - Exit 2: read the selected path's
--help; do not guess flags. - Exit 3: a recorded backend is unreachable; repair routing and reconfigure.
- Exit 4: run
vss configure --base-url <origin>or choose a path whose required services are actually routed. - Exit 5: ingest the source, wait for readiness, and re-run
vss configure. - Missing/ambiguous source: stop for clarification; never substitute.
- Missing RT-VLM: retrieval remains valid and results remain
unverified. - Authentication: use the operator-approved route. Never place secrets in prompts, flags, generated files, logs, or skill output.

