# Argus # Agents Argus exposes its documentation and two focused skills in formats an agent can fetch directly. Use `argus-setup` to operate a host and `argus-research` to turn an existing instance into sourced context without changing collection. ## Documentation interfaces [#documentation-interfaces] | Interface | Use | | -------------------------------------------------------- | --------------------------------------------------------------------------------------------- | | [`/llms.txt`](/llms.txt) | Compact index of every documentation page with absolute canonical URLs. | | [`/llms-full.txt`](/llms-full.txt) | Processed Markdown for the complete handbook. | | `/docs/.md` | Processed Markdown for one docs page, such as [`/docs/quick-start.md`](/docs/quick-start.md). | | [`/skill/SKILL.md`](/skill/SKILL.md) | The current portable `argus-setup` skill entry file. | | [`/skill/argus-skill.zip`](/skill/argus-skill.zip) | Deterministic archive containing `argus-setup`, its license, and references. | | [`/skill/research/SKILL.md`](/skill/research/SKILL.md) | The portable `argus-research` entry file. | | [`/skill/argus-research.zip`](/skill/argus-research.zip) | Deterministic research-skill archive with API, traversal, and provenance references. | ## Install the setup skill [#install-the-setup-skill] Download and extract the archive into your client's skills directory. Both Codex and Claude-compatible clients discover a directory containing `SKILL.md`; the archive already has the required `argus-setup/` top-level directory. ```bash curl -fsSLo /tmp/argus-skill.zip https://argus.gpsxtre.me/skill/argus-skill.zip unzip -o /tmp/argus-skill.zip -d ~/.codex/skills ``` For Claude Code, use its user skills directory instead: ```bash curl -fsSLo /tmp/argus-skill.zip https://argus.gpsxtre.me/skill/argus-skill.zip unzip -o /tmp/argus-skill.zip -d ~/.claude/skills ``` Restart or refresh the client if it does not discover newly installed skills. Inspect the entry file before installation when your client or organization requires a review step. ## Install the research skill [#install-the-research-skill] Install this alongside setup when an agent needs to query stored records, inspect X reply samples, traverse configured FxEmbed/SearXNG primitives, and write briefs with honest coverage language. ```bash curl -fsSLo /tmp/argus-research.zip https://argus.gpsxtre.me/skill/argus-research.zip unzip -o /tmp/argus-research.zip -d ~/.codex/skills ``` Use `~/.claude/skills` instead for Claude Code. API tokens belong in the agent's runtime environment. The research skill never instructs an agent to place a bearer token in a URL, transcript, or report. ## CLI and safety boundaries [#cli-and-safety-boundaries] Use `--json` whenever an agent reads CLI output. Successful responses use the versioned envelope `{"contractVersion":1,"ok":true,"data":{}}`; failures contain `error.code`, `error.message`, and sometimes `error.recovery`. Treat the live CLI schema, help, plan, and JSON result as authority rather than recreating deployment logic in a chat response. Read-only inspection is distinct from change. Before onboarding, repair, update, configuration apply, or any operation that can change a host or an optional Cloudflare account, inspect the CLI plan, explain it, and obtain explicit approval at the CLI confirmation boundary. A dry run does not authorize its later mutating form. Secrets must stay in hidden CLI prompts or approved local secret storage, never in chat, an answers file, a command argument, or a transcript. The CLI redacts known secret values in its output, but agents must still avoid echoing, persisting, or forwarding credentials. For a failed health check or command, run `argus doctor --json`, report its component/code/message, and follow a returned recovery or log command exactly. If recovery is missing, unclear, outside approval, or requests new authority, stop and ask the user. Do not edit Compose files, deployment state, backups, managed settings, or secret files directly; the [setup skill](/skill/SKILL.md) contains the detailed routing rules. # API reference `GET /health` is public. Every `/v1/*` route requires `Authorization: Bearer $ARGUS_API_TOKEN` when a token is configured. Consumers use their own bearer credential for records, watches, collections, changes, usage, source health, and scoped configuration export. Other routes require the operator token. Configuring consumers requires an operator token, including on loopback. Primitive routes require a configured operator token. ```bash export ARGUS_URL=http://127.0.0.1:8788 curl -H "Authorization: Bearer $ARGUS_API_TOKEN" "$ARGUS_URL/v1/records" ``` ## `GET /health` [#get-health] `GET /health` returns `{status, version: 2, role, intelligence}`. ## `GET /v1/records` [#get-v1records] `GET /v1/records` accepts repeated `source` and `target`, plus `q`, `since`, `until`, `publishedSince`, `publishedUntil`, `collectedSince`, `collectedUntil`, `limit` (1–200, default 50), and `cursor`. Pass `nextCursor` back unchanged. `since`/`until` filter collection checks; publication filters independently select source publication times. This endpoint reads stored evidence and performs no fresh collection. Use `/v1/changes` for reliable delivery; record-list pagination is not a change stream. Each compact envelope includes `id`, `revisionId`, `source`, `url`, optional author and publication time, `firstCollectedAt`, `lastCheckedAt`, `observedChangedAt`, a 500-character `excerpt`, `truncated`, `detail`, and available `related` content. Unknown publication times remain absent. `observedChangedAt` is Argus's observed change time, not a claimed source edit time. ## `GET /v1/records/:id` [#get-v1recordsid] `GET /v1/records/:id` returns public evidence, media pointers, relations, and latest engagement; private watch configuration and provider raw payloads are excluded. It returns 404 when absent. `offset` and `relatedOffset` default to 0. Text pages contain at most 16,384 characters; media and relation pages each contain at most 50 entries. Follow `pagination.nextOffset` and `pagination.nextRelatedOffset`. The response reports total lengths and `pagination.truncated`. ## `GET /v1/records/:id/conversation` [#get-v1recordsidconversation] `GET /v1/records/:id/conversation` returns the root record, the latest ranked X reply sample with hydrated reply records, and bounded snapshot history. History accepts `limit` (1–100, default 20) and an opaque `cursor`. Fetch each reply through the detail route. ```bash curl -H "Authorization: Bearer $ARGUS_API_TOKEN" \ "$ARGUS_URL/v1/records?source=x&since=2026-08-29T00%3A00%3A00.000Z&limit=50" curl -H "Authorization: Bearer $ARGUS_API_TOKEN" \ "$ARGUS_URL/v1/records/$RECORD_ID/conversation" ``` ## `GET /v1/artifacts` [#get-v1artifacts] `GET /v1/artifacts?kind=summary&limit=20` returns derived artifacts with exact record/media provenance. ## `POST /v1/summaries` [#post-v1summaries] `POST /v1/summaries` accepts `query`, `watchIds`, `limit`, and `prompt`. It requires enabled OpenRouter intelligence and stores a `summary` artifact. ## `POST /v1/query` [#post-v1query] `POST /v1/query` accepts a required `question` and optional `watchIds`, `since`, `until`, and `limit` (1–100). It selects recent matching context, produces a sourced OpenRouter answer, stores an `answer` artifact, and returns numbered source links. Both summaries and answers include each selected root's latest bounded reply sample when available, and store the exact snapshot/reply IDs in artifact provenance. The CLI command is the easiest client: ```bash argus query latest-news --watch=screen-news --since=2026-08-29T09:00:00.000Z --until=2026-08-29T18:00:00.000Z --limit=50 --json curl -X POST -H "Authorization: Bearer $ARGUS_API_TOKEN" \ -H 'content-type: application/json' \ --data '{"question":"new listings since 9am","watchIds":["listings"],"since":"2026-08-29T09:00:00.000Z"}' \ "$ARGUS_URL/v1/query" ``` ## `POST /v1/watches/:watchId/ingest` [#post-v1watcheswatchidingest] `POST /v1/watches/:watchId/ingest` queues each target in an enabled watch and returns 202 with `{queued, watchId, requests}`. Consumers may ingest only their own watches; each request is subject to collection quotas. ## `GET /v1/primitives/x/*` [#get-v1primitivesx] `GET /v1/primitives/x/2/*` proxies a normalized `/2/` path to the configured FxEmbed origin. Query parameters pass through. The agent cannot provide a host. Only GET and HEAD are supported; response bodies are capped at 2 MiB, same-origin redirects at five, requests at ten seconds, and each token/source at 60 calls per minute. Every public DNS answer is validated and pinned again for every redirect hop. ## `HEAD /v1/primitives/x/*` [#head-v1primitivesx] HEAD uses the same fixed-origin, path, authentication, redirect, timeout, and rate limits as GET, but returns no response body. ```bash curl -H "Authorization: Bearer $ARGUS_API_TOKEN" \ "$ARGUS_URL/v1/primitives/x/2/conversation/1900?cursor=next" ``` Primitive output is transient: Argus does not write records, checkpoints, artifacts, or jobs for this request. ## `GET /v1/primitives/web/search` [#get-v1primitiveswebsearch] `GET /v1/primitives/web/search?q=...` calls only the configured SearXNG service and forces `format=json`. Public endpoints receive DNS/private-address validation on every hop; only endpoints explicitly marked `searchEndpointTrust: trusted` may use a private operator-controlled network. Optional allowlisted parameters are `engines`, `categories`, `language`, `time_range`, and `pageno`; every other parameter is rejected. It uses the same authentication, body, redirect, timeout, and rate bounds and also writes nothing. ```bash curl -H "Authorization: Bearer $ARGUS_API_TOKEN" \ "$ARGUS_URL/v1/primitives/web/search?q=new+movies" ``` ## `POST /v1/diagnostics/smoke-watches` [#post-v1diagnosticssmoke-watches] Creates a temporary watch for `argus doctor`. Its records are isolated from normal watches. ## `GET /v1/diagnostics/smoke-watches/:id/records` [#get-v1diagnosticssmoke-watchesidrecords] Returns records collected for one temporary smoke watch. ## `DELETE /v1/diagnostics/smoke-watches/:id` [#delete-v1diagnosticssmoke-watchesid] Removes the temporary watch and its isolated records. ## `POST /v1/management/config/plan` [#post-v1managementconfigplan] Previews the managed configuration operations used by the CLI. The result includes `planId`, `baseContentHash`, `desiredContentHash`, and `removals` with each removed watch ID and optional owner. ## `POST /v1/management/config/apply` [#post-v1managementconfigapply] Applies a previously planned managed configuration change. The exact preview is required; a changed applied snapshot or edited preview returns 409. ## `POST /v1/management/config/verify` [#post-v1managementconfigverify] Verifies the resulting managed deployment. Management routes require an exact configured bearer token; prefer the CLI rather than calling them directly. ## Consumer watches and configuration [#consumer-watches-and-configuration] | Endpoint | Behavior | | ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | | `GET /v1/watches` | List your watches, including disabled watches; `limit` 1–100 (default 50), `offset` default 0, and `nextOffset` for more. | | `POST /v1/watches` | Create a watch; return 201. | | `PUT /v1/watches/:id` | Replace a watch using its complete configuration. | | `DELETE /v1/watches/:id` | Disable future collection; preserve evidence. | | `GET /v1/config/export` | Export your watches and quotas without credentials. | | `POST /v1/management/config/export` | Operator-only export of the applied configuration. | Watch bodies use the [configuration schema](/docs/configuration): `id`, `schedule`, `inputs`, optional `enabled` (default true), and `classify`. Ownership is assigned to the authenticated consumer. A consumer cannot transfer ownership, read another consumer's watch, or mutate it. Duplicate IDs and concurrent configuration changes return 409; exhausted watch capacity returns 429 with `error: "watch_quota"`. `argus config export` includes agent-created watches in the same applied snapshot used by operator previews. Exported credentials must be restored from secret references before reapplying. ## Fresh collection [#fresh-collection] ### `POST /v1/collections` [#post-v1collections] `POST /v1/collections` returns a durable request immediately (202), or a throttled request (429). JSON request bodies are capped at 65,536 bytes. ```json { "source": "x", "target": { "kind": "post", "value": "1900000000000000000" }, "coverage": { "linkedArticles": 0, "replies": 20, "media": false }, "maxAgeSeconds": 60, "priority": "interactive" } ``` | Field | Accepted values and defaults | | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `source` | `x`, `telegram`, or `web`; the source must be enabled. | | `target.kind` | X: `post`, `account`, `query`; Telegram: `channel`; web: `url`, `feed`, `query`. | | `target.value` | Nonempty string, at most 4,096 characters. X posts accept IDs or public X/Twitter post URLs. Accounts/channels omit `@`. Web URLs must be HTTP(S) without credentials. | | `coverage.linkedArticles` | 0–10, default 0; follows available linked article relations. | | `coverage.replies` | 0–200, default 0; bounded X reply sampling. | | `coverage.media` | Boolean, default false; available media metadata only. Binary extraction is unsupported. | | `maxAgeSeconds` | 0–86,400, default 0; acceptable reuse age. | | `priority` | `interactive` (default) or `background`. | Root evidence is always requested. Root URL/post collection is bounded to one root; account, channel, feed, and query collection to 100 source items. Requested expansion can remain missing: inspect the receipt instead of assuming complete coverage. Background jobs receive scheduling turns alongside interactive work. Equivalent requests share upstream work when target, coverage, and freshness requirements permit; each consumer retains its own request and cancellation. | Endpoint | Behavior | | --------------------------------- | ------------------------------------------------------------------------ | | `GET /v1/collections` | Your requests; `limit` 1–100 (default 50), opaque `cursor`. | | `GET /v1/collections/:id` | Request status and available partial receipt. | | `GET /v1/collections/:id/receipt` | Receipt, or 202 with `receipt: null` while unavailable. | | `PATCH /v1/collections/:id` | Change a queued request's `priority`; returns 409 when no longer queued. | | `DELETE /v1/collections/:id` | Cancel your unsettled request; other consumers' shared work continues. | Request status is `queued`, `running`, `complete`, `failed`, `cancelled`, or `throttled`. Requests expose stable `id`, `workId` when attached, `reused`, `remainingRequests`, and optional `throttleReason`: `request_quota`, `active_jobs`, or `minimum_refresh`. Receipts identify checked sources, start/completion times, requested and obtained coverage, missing coverage, truncation, failures, and exact `recordId`/`revisionId` pairs. Receipt revisions are paginated in groups of 100; pass `pagination.nextOffset` as `offset` to the receipt endpoint. Request lists include `receiptUrl` rather than embedded receipts. Partial evidence remains reusable after failures. Default consumer limits are 100 enabled watches, 60 requests per hour, four active jobs, 60 upstream requests per hour, and a 60-second minimum refresh interval. Configure these independently in [consumer configuration](/docs/configuration). Reuse counts as a consumer request, without counting as another upstream fetch. ## Immutable revisions [#immutable-revisions] ### `GET /v1/records/:id/revisions` [#get-v1recordsidrevisions] `GET /v1/records/:id/revisions` lists revision metadata using `limit` 1–100 (default 50) and an opaque `cursor`. ### `GET /v1/records/:id/revisions/:revisionId` [#get-v1recordsidrevisionsrevisionid] This endpoint retrieves that exact revision with the same `offset`/`relatedOffset` bounds as record detail. Unknown or expired revisions return 410 with `error: "REVISION_UNAVAILABLE"`; current content is never substituted for an unavailable historical revision. ## Durable changes and resume [#durable-changes-and-resume] ### `GET /v1/changes` [#get-v1changes] `GET /v1/changes` accepts an opaque `cursor`, `limit` 1–200 (default 50), and repeated `type`, `source`, and `recordId` filters (at most 100 record IDs). Without an explicit cursor it uses your saved delivery position. ### `PUT /v1/changes/cursor` [#put-v1changescursor] Save progress only after processing the page using `{"cursor":"..."}`. `cursor=latest` obtains a starting position for future changes. Event types are `record.created`, `record.revised`, `reply.sampled`, `engagement.updated`, `record.deleted`, `job.completed`, and `source.health`. Delivery ordering is independent of publication time. Use stable event IDs and record/revision IDs to handle repeated pages safely. Unchanged checks do not create content revision events, and failed fetches do not imply deletion. An expired cursor returns 410 with `error: "CHANGE_CURSOR_EXPIRED"` and explicit resynchronization links. Save a new latest cursor first, scan stored records, then resume from that cursor; deduplicate any overlap. Event retention defaults to 30 days; record and revision retention each default to 90 days. ## Usage and health [#usage-and-health] ### `GET /v1/usage` [#get-v1usage] `GET /v1/usage` reports your requests, reuse, throttling, upstream fetches, active jobs, and request quotas. Optional `since` is a UTC timestamp; the default window is the previous hour. ### `GET /v1/sources/health` [#get-v1sourceshealth] This endpoint accepts `limit` 1–100 (default 50) and an opaque `cursor`. It reports an opaque target ID, source status, check time, successful-check time when known, consecutive failures, and `lagSeconds` since success (null when unknown). Private target configuration and provider errors are excluded. Health transitions are delivered through `source.health` events, separating a quiet source from an unavailable one. Collection, revision retrieval, delivery, and health require no LLM calls. `collectedSince` and `collectedUntil` filter first collection time; `since` and `until` filter the last check. Scalar fields shortened for response limits report `truncatedFields`; text and related arrays use explicit pagination. Watch lists accept `offset` and `limit` (1–100). `/v1/usage` includes remaining request, upstream, and active-job capacity for the current hour. # Core concepts Argus is a context data layer. Deterministic ingestion keeps working without an LLM; optional OpenRouter intelligence reads the stored evidence later. ## From watch to record [#from-watch-to-record] A **watch** names a schedule and a set of X, public Telegram, and Web inputs. Each input expands into a deterministic **target**. The scheduler persists due jobs, a worker leases them, and one commit writes the normalized records plus the target checkpoint. A canonical record ID is the SHA-256 identity of `source + externalId`. It does not contain a watch or target, so the same source item observed by two watches is stored once and linked to both observations. Records preserve source URL, text, raw payload, first/last-seen times, media pointers, relations, and latest engagement. Changed content creates an immutable revision; engagement-only changes create engagement snapshots instead. ## Media pointers, never media bytes [#media-pointers-never-media-bytes] Argus stores public URLs and metadata for images, video, audio, and documents. It never mirrors the bytes. This keeps the deterministic data layer small and lets a downstream agent decide whether and how to retrieve a source asset. ## X conversations [#x-conversations] Reply tracking is separate work from root-post ingestion. A successful root record can start an independent conversation tracker. Refresh jobs follow distinct upstream cursors until exhaustion or 500 distinct replies, select the configured bounded sample, store those replies as ordinary canonical records, and atomically save their ranks in a conversation snapshot. “50 most liked” means the 50 most-liked replies Argus observed, not a claim about every reply on X. Each snapshot records observed count, retained count, ordering, page count, completeness, truncation reason, and collection time. ## Artifacts and provenance [#artifacts-and-provenance] Summaries and query answers are **derived artifacts**, never replacements for records. An artifact points to the exact record and media IDs used, provider, model, generation ID, prompt/question, and media dispositions. Social reply evidence is marked as an observed bounded sample. ## Transient primitives [#transient-primitives] Authenticated agents can use Argus as a bounded proxy for configured FxEmbed v2 GET/HEAD calls and SearXNG search. Primitive responses are not persisted. The caller cannot choose an upstream origin; paths, redirects, time, response size, and rate are bounded. This enables traversal without turning every intermediate request into permanent data. ## SQLite and PostgreSQL [#sqlite-and-postgresql] SQLite runs the full stack in one process and is the easiest VPS deployment. PostgreSQL implements the same storage contract and permits separate runtime roles. Argus v2 refuses legacy database schemas before mutation instead of silently mixing incompatible data models. # Configuration Argus reads one strict YAML file. Unknown fields and `version: 1` are rejected; v2 deliberately starts from a fresh database. Put secrets in `/opt/argus/secrets.env` and reference them as a complete `${NAME}` value. ```bash argus config validate /opt/argus/argus.yaml argus config show /opt/argus/argus.yaml argus config apply /opt/argus/argus.yaml --dry-run argus config apply /opt/argus/argus.yaml --yes argus config schema --json ``` ## Complete v2 examples [#complete-v2-examples] ### Single VPS with SQLite [#single-vps-with-sqlite] ```yaml config version: 2 runtime: { role: all } storage: adapter: sqlite url: /app/data/argus.db sources: x: enabled: true endpoint: http://fxembed:8787 replies: enabled: true maxPerPost: 50 maxTrackingHours: 168 orderBy: likes telegram: { enabled: true, adapter: public-web } web: enabled: true searchEndpoint: http://searxng:8080 searchEndpointTrust: trusted userAgent: Argus/0.1 watches: - id: screen-news enabled: true schedule: "*/10 * * * *" inputs: x: accounts: [DiscussingFilm, FilmUpdates] queries: ["new movies and TV shows"] telegram: { channels: [mcunewsandrumors] } web: urls: [https://deadline.com/v/tv/] feeds: [https://deadline.com/feed/] queries: ["new movies and TV shows news"] classify: keywords: [movie, film, tv, series, trailer, streaming] api: host: 0.0.0.0 port: 8788 token: ${ARGUS_API_TOKEN} ``` ## Storage [#storage] `storage.adapter` is `sqlite` or `postgres`. * SQLite is the one-command VPS default and requires `runtime.role: all`. * PostgreSQL accepts a canonical `postgres://` or `postgresql://` URL and can support separate `api`, `scheduler`, `worker`, and `processor` processes. * A v1 or unversioned database is refused before mutation. Argus does not run a compatibility migration; start v2 with an empty database. Both adapters implement the same record, revision, media, relation, engagement, conversation, job, checkpoint, artifact, and configuration contracts. ## X reply tracking [#x-reply-tracking] `sources.x.endpoint` is the selected FxEmbed origin; VPS onboarding writes `http://fxembed:8787`. Reply tracking is opt-in in hand-written YAML. The onboarding wizard recommends the Standard profile. | Profile | `maxTrackingHours` | Intended use | | -------- | -----------------: | -------------------------------------- | | Hot | 24 | Fast-moving, high-volume topics | | Standard | 168 | General monitoring; onboarding default | | Niche | 720 | Slow-moving specialist topics | `maxPerPost` defaults to 50 and is bounded from 1 to 200. Argus retains the top configured number from the replies it actually observed; it never claims the sample is every reply on X. `orderBy` accepts `likes`, `newest`, `oldest`, `replies`, `reposts`, `views`, or `source`. Growth temporarily increases refresh frequency. Tracking stops at the configured time horizon even if the retained sample has not reached 50. ## Telegram and Web [#telegram-and-web] Telegram v2 supports public announcement channels through `adapter: public-web`. No bot token, private chat, member-only group, or generic discussion ingestion is implemented. Direct Web URLs and feeds use the bounded public-Web fetch path. Search queries require `searchEndpoint`, normally managed SearXNG. External endpoints use `searchEndpointTrust: public`; only an operator-controlled private service uses `trusted`. ## Watches [#watches] A watch needs a lowercase `id`, a five-field cron `schedule`, and `inputs`. Inputs may contain X `accounts` and `queries`, Telegram `channels`, and Web `urls`, `feeds`, and `queries`. `classify.keywords` adds matching metadata; it does not discard unmatched records. ## Intelligence [#intelligence] Intelligence is optional and OpenRouter-only. Scheduled processors create sourced summary artifacts. `argus query` creates a sourced answer artifact on demand. Supported public image, video, and PDF pointers may be passed directly to a compatible model; Argus never downloads or base64-encodes media. ## API and secrets [#api-and-secrets] `api.token` protects every `/v1/*` endpoint. The runtime refuses a non-loopback bind without it. The managed Compose topology publishes the selected port, so also use a firewall, private network, or reverse proxy policy. Use `argus secrets set ARGUS_API_TOKEN` and `argus secrets set OPENROUTER_API_KEY`; never place their values in YAML or a shell argument. ### PostgreSQL service roles [#postgresql-service-roles] ```yaml config version: 2 runtime: { role: all } storage: adapter: postgres url: postgresql://postgres:5432/argus?sslmode=disable sources: {} watches: [] intelligence: enabled: true provider: openrouter apiKey: ${OPENROUTER_API_KEY} model: openai/gpt-4.1-mini processors: [] api: host: 0.0.0.0 port: 8788 token: ${ARGUS_API_TOKEN} ``` ### Minimal Web-only instance [#minimal-web-only-instance] ```yaml config version: 2 runtime: { role: all } storage: { adapter: sqlite, url: /app/data/argus.db } sources: web: { enabled: true } watches: - id: public-site schedule: "*/15 * * * *" inputs: web: { urls: [https://example.com/news] } api: host: 0.0.0.0 port: 8788 token: ${ARGUS_API_TOKEN} ``` ## Consumers and evidence retention [#consumers-and-evidence-retention] `consumers` defaults to an empty list. Each consumer has a unique `id` and its own `credential` (a literal token or `${ARGUS_CONSUMER_TOKEN}` secret reference). A consumer without a credential cannot authenticate. Configuring any consumers requires an operator `api.token`, including on loopback. Consumer credentials must be distinct from one another and from the operator's `api.token`. Per-consumer `quotas` default to `maxWatches: 100`, `maxJobsPerHour: 60`, `maxConcurrentJobs: 4`, and `maxUpstreamRequestsPerHour: 60`. `minimumRefreshSeconds` defaults to 60. Zero quota disables that capacity. An owned watch sets `ownerId` to its consumer's ID; other consumers cannot mutate it. `retention.recordsDays` and `retention.revisionsDays` each default to 90; `retention.eventsDays` defaults to 30. All retention periods are positive integers. `argus config export` exports the service's applied configuration, including agent-created watches, with credentials removed. Restore secret references before applying an exported configuration. `argus config show` displays the local file. Configuration apply previews removed watch IDs and owners. The preview's plan ID and base hash bind approval to the current applied configuration; changes after preview invalidate the plan and require a new preview. Cancelling a watch preserves its collected evidence. # Deployment ## Recommended topology: one VPS and SQLite [#recommended-topology-one-vps-and-sqlite] Run one Argus instance on a supported Linux VPS. This is the supported onboarding path and the smallest complete deployment: one Argus service runs the API, scheduler, worker, and optional processor with `runtime.role: all` and SQLite. Use one of Ubuntu 22.04, Ubuntu 24.04, Ubuntu 25.10, Ubuntu 26.04, Debian 12, or Debian 13 on Linux amd64/x64 or arm64/aarch64. The host needs a sudo-capable account, outbound HTTPS, Docker Engine with the Compose plugin and daemon access, at least 5 GiB free disk, and 1 GiB memory. Choose 2 GiB when onboarding managed SearXNG. Onboarding also verifies that the selected API port is free before it writes an instance. Inspect the installer before making host changes, then install and onboard: ```bash curl -fsSLo /tmp/argus-install.sh https://argus.gpsxtre.me/install.sh ARGUS_INSTALL_INSPECT=1 sh /tmp/argus-install.sh # After review, run the exact file inspected above. sh /tmp/argus-install.sh argus onboard argus status --json argus doctor --json ``` The instance root is `/opt/argus`. The CLI owns its generated Compose file and state; do not edit them to change a deployment. ## What the generated Compose deployment contains [#what-the-generated-compose-deployment-contains] Every managed Compose deployment has these networks: * `argus-private` is internal. Argus, PostgreSQL when selected, managed SearXNG, and VPS-hosted FxEmbed use it. * `argus-egress` lets source-facing services reach the public Internet. Argus, managed SearXNG, and VPS-hosted FxEmbed use it. The `argus` service mounts `/opt/argus/argus.yaml` read-only and the persistent `argus-data` volume at `/app/data`. It reads `/opt/argus/secrets.env` as its environment file. Docker publishes the selected API port, default `8788`, from the host to the container. Put a firewall or reverse proxy in front of a publicly reachable host port; setting a loopback API host is not the same thing as removing Docker's published port. SQLite is stored through `argus-data`; there is no separate database container. Selecting PostgreSQL adds a `postgres` service and the persistent `postgres-data` volume. Selecting managed SearXNG adds a `searxng` service with the CLI-owned `/opt/argus/searxng/settings.yml`. Selecting VPS FxEmbed adds the signed, digest-pinned `fxembed` service. Neither service publishes a host port. `state.json` records the verified deployment and pinned image references. `release-context.json` keeps the signed release material used to identify a rollback release. `backups/` and `update-state.json` are update recovery state. They are operational state, not application configuration; preserve them and do not hand-edit them. ## Managed and external dependencies [#managed-and-external-dependencies] Managed SearXNG is available only for Web query targets. Argus checks it with a bounded JSON search and can recreate only this managed service with `argus repair searxng`. An external SearXNG endpoint is outside Argus's control: Argus can check it, but it will not repair or change it. VPS-hosted FxEmbed is the default for X. Argus runs the pinned Worker bundle through its local Workers runtime at `http://fxembed:8787`, reachable only inside Compose. It participates in lifecycle, status, logs, doctor, targeted repair, signed update, and rollback. Cloudflare-hosted and external FxEmbed remain advanced alternatives; Argus cannot repair an external endpoint or Cloudflare account configuration. ## Advanced PostgreSQL roles [#advanced-postgresql-roles] Use PostgreSQL when the runtime needs independent `api`, `scheduler`, `worker`, or `processor` processes. SQLite is invalid for every role except `all`; it is not a multi-process or multi-role storage mode. All role processes use the same validated configuration and PostgreSQL database. `ARGUS_ROLE` overrides `runtime.role` for a process. Run at least an API role for the HTTP API, a scheduler role to enqueue due work, and a worker role to collect it; add a processor role only for scheduled intelligence processors. PostgreSQL job leases expire, so a crashed worker's work can be retried without Redis or Kafka. The generated VPS Compose topology still starts a single `all` Argus service. Splitting roles is an advanced operator-managed deployment concern; the CLI does not render a multi-replica Compose layout. ## Railway boundary [#railway-boundary] The repository includes Railway templates in `deploy/railway/` for `api`, `scheduler`, `worker`, and `processor` roles. Use PostgreSQL and set `ARGUS_ROLE` for each service. These templates are not a Railway control plane: `argus onboard`, lifecycle commands, managed Compose repair, update backup, and rollback operate on the `/opt/argus` Docker VPS instance, not on Railway services. Use Railway's own deployment, secret, database backup, and recovery controls for that topology. ## Backup boundary [#backup-boundary] The managed SQLite database is in Docker's named `argus-data` volume, which Compose creates as `argus_argus-data` because the project name is `argus`. Confirm that identity without changing the instance before every backup or update: ```bash cd /opt/argus docker compose -p argus config --volumes docker volume inspect argus_argus-data ``` For SQLite, a signed update proves that exact Compose-owned volume, stops the Argus writer, and creates a consistent snapshot beneath `/opt/argus/backups`. The updater records its hash, size, integrity result, table counts, and volume identity before pulling or migrating. `argus update --rollback` revalidates the snapshot and atomically restores it inside `argus_argus-data` before starting the prior signed release. Keep a separately verified operator-managed Docker-volume backup for disaster recovery; the update snapshot is not a substitute for an independent backup policy. For PostgreSQL, take regular `pg_dump`/`pg_restore` backups or provider snapshots and store them outside `/opt/argus`. Argus deployment backups do not contain a PostgreSQL dump. After a database restore, start the required runtime roles and verify with `argus status --json` and `argus doctor --json` on a managed VPS. ## Next step [#next-step] Read [operations](/docs/operations) for inspect/apply/verify sequences and [security](/docs/security) before exposing the API or granting optional Cloudflare credentials. # Argus documentation Argus continuously collects public signals, stores canonical records and their revisions, and exposes deterministic queries with source links. Optional summaries are derived artifacts; ingestion and storage do not depend on an LLM. ## Choose your path [#choose-your-path] * New operator: [run the quick start](/docs/quick-start). * Existing operator: [open operations](/docs/operations) or [troubleshooting](/docs/troubleshooting). * Agent: use the [machine-readable interfaces](/docs/agents). * Contributor: start with the [contributor guide](/docs/contributing). ## Supported in v1 [#supported-in-v1] | Area | Current capability | | --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | | Public sources | X accounts and queries through FxEmbed, public Telegram announcements, and Web URLs, feeds, and SearXNG-backed queries. | | Storage | SQLite for a single all-in-one runtime, or PostgreSQL for shared multi-role deployments. | | Scheduling and query | Cron-derived target jobs, leased workers, retries and checkpoints, plus authenticated deterministic record and artifact queries. | | Optional intelligence | OpenRouter-backed sourced summaries stored separately as artifacts. | | Deployment | Signed release installation on a supported VPS Docker host, with Argus-managed Docker Compose and private, optional SearXNG and FxEmbed services. | | Runtime boundaries | API, scheduler, worker, and processor responsibilities can run together; SQLite requires the combined role while PostgreSQL supports separated roles. | ## Product boundary [#product-boundary] Argus does not access private Telegram conversations, bypass site controls, or require an LLM for deterministic ingestion and querying. # Installation and onboarding ## Outcome [#outcome] This procedure installs the signed `argus` host wrapper and creates a verified, managed Argus instance at `/opt/argus`. It is for a fresh VPS Docker host, not a workstation and not a manually maintained existing instance. ## Prerequisites [#prerequisites] The clean-host installer matrix covers Ubuntu 22.04 and 24.04 plus Debian 12 and 13, across Linux AMD64 and ARM64 where represented by the matrix. Release images are built for `linux/amd64` and `linux/arm64`. Use `x86_64`/`amd64` or `aarch64`/`arm64`, a sudo-capable account, outbound HTTPS, and at least 5 GiB free disk space. Onboarding checks Docker Engine, Docker Compose, daemon access, memory (at least 1 GiB, or 2 GiB with managed SearXNG), and the selected API port before it changes the instance. ## Inspect the installer before it changes a host [#inspect-the-installer-before-it-changes-a-host] Download the installer from the public distribution URL, inspect it, then run its no-change inspection mode: ```bash curl -fsSLo /tmp/argus-install.sh https://argus.gpsxtre.me/install.sh less /tmp/argus-install.sh ARGUS_INSTALL_INSPECT=1 sh /tmp/argus-install.sh ``` The generated installer embeds the Ed25519 trust root. During installation it downloads `manifest.json` and `manifest.sig`, verifies the manifest signature before trusting its contents, validates the canonical manifest shape, then verifies the downloaded wrapper's SHA-256 hash. A failed signature, malformed manifest, or hash mismatch stops before the wrapper is replaced. Install only after reviewing the report: ```bash curl -fsSL https://argus.gpsxtre.me/install.sh | sh ``` ## Durable management launcher [#durable-management-launcher] The signed installer writes the strict `/opt/argus/management.state` file and installs one immutable `/usr/local/bin/argus` launcher. Later `argus update` commands advance the verified management state without replacing the launcher. The next release refuses legacy version-pinned wrappers; remove an obsolete launcher explicitly before installing it. Do not edit `management.state`. A missing or malformed state makes the launcher fail closed before Docker starts. Rerun the signed installer to repair the launcher/state pair. ## Docker-present and Docker-absent paths [#docker-present-and-docker-absent-paths] With usable Docker Engine and the Compose plugin, the installer downloads and installs the verified wrapper. With Docker absent, an interactive run asks for approval before it installs Docker from Docker's official apt repository. For a non-interactive approved install, download the script first and run: ```bash ARGUS_INSTALL_DOCKER=1 sh /tmp/argus-install.sh ``` Use `ARGUS_INSTALL_DOCKER=0` when Docker installation is forbidden; in that case the installer reports a missing Docker dependency instead of installing it. It also refuses conflicting distribution Docker packages and stops if the daemon or Compose plugin is unavailable. ## Interactive onboarding answers [#interactive-onboarding-answers] Run interactive onboarding after the wrapper is installed: ```bash argus onboard ``` The answer groups appear in this order: * **Storage and API** always appear: SQLite (the default) or PostgreSQL, then the API port (default `8788`). * **Sources and watch** always appear: choose X, public Telegram announcements, and/or Web; then give the watch ID and cron schedule. The selected source receives only its relevant inputs: X accounts and queries, public Telegram channel names, or Web URLs, feeds, and queries. * **SearXNG** appears only when at least one Web search query is entered. Choose a managed service or an external endpoint. * **FxEmbed** appears only when X is selected. The default runs FxEmbed on the same VPS. It stays on Argus's private Compose network and publishes no host port. Cloudflare deployment and an external endpoint are advanced choices. * **Intelligence** always appears as the OpenRouter summaries choice. When enabled, it asks for the model; when disabled, no OpenRouter secret is needed. * **Keywords** always appear and become watch classification metadata. * **Secrets** are hidden prompts: an Argus API token always; a PostgreSQL password for PostgreSQL storage; a Cloudflare API token only when Cloudflare FxEmbed is explicitly selected; and an OpenRouter API key for intelligence. Telegram is limited to public channels. Do not enter private conversations or private channels. Managed SearXNG is meaningful only for Web query targets; direct URLs and feeds do not require it. ## File-driven onboarding and secret handling [#file-driven-onboarding-and-secret-handling] For repeatable non-secret choices, use a strict YAML answers file. It must use version `1`, require `/opt/argus` as its root, match the current onboarding schema, and not be group- or world-writable. It must never include a secret field. Inspect the plan without mutation with: ```bash argus onboard --from answers.yaml --dry-run --json ``` After reviewing that plan, run the file-driven apply from an interactive terminal so the CLI can collect any required hidden secrets. Use `--yes` only when you have already reviewed the current plan. ```bash argus onboard --from answers.yaml ``` The generated secret file is `/opt/argus/secrets.env`. It is a regular, owner-owned `0600` file; Argus rejects symlinks, non-files, and unsafe modes. Secrets are referenced from generated configuration but are not written into the versioned YAML answers file or the applied configuration snapshot. ## Managed files and safe repeatability [#managed-files-and-safe-repeatability] Onboarding manages these instance files: * `/usr/local/bin/argus` — the verified host wrapper installed by the signed installer. * `/opt/argus/argus.yaml` — generated runtime configuration, written after validation. * `/opt/argus/secrets.env` — owner-only secret environment file. * `/opt/argus/compose.yaml` — Argus-managed Docker Compose definition. * `/opt/argus/state.json` — validated persisted deployment state. * `/opt/argus/searxng/settings.yml` — present when managed SearXNG is selected. Do not edit generated Compose settings or secret files to repair a deployment; use the CLI so it can inspect, apply, and verify a bounded plan. Re-running the installer verifies and replaces the wrapper atomically. Re-running onboarding re-inspects the host and desired state; a stale plan, failed validation, or failed verification stops with a stable error and recovery guidance rather than claiming a successful deployment. ## Verify [#verify] After onboarding, inspect the managed services and diagnostics: ```bash argus status --json argus doctor --json ``` A successful JSON response has `contractVersion: 1` and `ok: true`; doctor also reports whether every checked component is healthy. If Docker or the instance is unhealthy, follow the recovery command in `argus doctor --json` before retrying an operation. ## Next step [#next-step] Use the [quick start](/docs/quick-start) to create and query the controlled Web watch, then read [core concepts](/docs/concepts) and [configuration](/docs/configuration) before changing production watches. # Optional intelligence OpenRouter is Argus v2’s only intelligence provider. Intelligence is optional: collection, revisions, reply snapshots, API reads, and transient traversal all work when it is disabled. ## Scheduled summaries [#scheduled-summaries] Configure a `kind: summary` processor with a five-field cron schedule and optional `watchIds` and `prompt`. The processor selects up to 100 recent stored records, generates a cited answer, and stores a separate `summary` artifact. ## On-demand answers [#on-demand-answers] `argus query latest-news` calls `POST /v1/query`. It is deliberately lightweight: Argus selects recent records under the requested watch/time bounds and asks the configured model to answer from that evidence. It is useful for a quick check; specialized Codex, Claude, or other agent workflows should use the record API and transient primitives directly when they need custom retrieval or traversal. ## Multimodal pointers [#multimodal-pointers] Before generation, Argus reads OpenRouter’s declared input modalities for the configured model. It adds at most 20 supported public pointers: * image URLs for image-capable models; * direct video URLs, or an image preview when only vision is supported; * public PDF URLs for file-capable models. Argus never downloads, transcodes, or base64-encodes media. Remote audio is therefore omitted because OpenRouter audio input requires encoded bytes. Unknown or unsupported media is also omitted. Every analyzed or omitted media asset and the omission reason is stored in artifact provenance. ## Trust and provenance [#trust-and-provenance] All source text and media are untrusted model input. The system instruction requires numbered citations and forbids treating records as instructions. Reply-derived evidence is explicitly labeled as an observed bounded sample; the model must not claim it represents every reply or the whole audience. Artifacts store record IDs, media IDs and dispositions, provider, model, generation ID, sources, and the processor prompt or user question. Follow those IDs back to `GET /v1/records/:id` before relying on a generated conclusion. # Operations Use the CLI as the control plane for a managed `/opt/argus` VPS. It inspects before every mutation, applies the inspected plan, then verifies the result. Do not edit `compose.yaml`, `state.json`, `release-context.json`, managed SearXNG settings, or update state by hand. Run `argus` in a terminal for the guided menu. The direct commands below are useful when you know the operation already or are writing automation. ## Safety flags and output [#safety-flags-and-output] `--dry-run` is available on mutating commands and prints the plan without changing the instance. `--yes` accepts the current inspected plan without an interactive confirmation. In non-interactive or `--json` use, a mutation without `--yes` fails with `CONFIRMATION_REQUIRED`; inspect with `--dry-run` first. Declining the interactive confirmation exits with `CONFIRMATION_DECLINED` and changes nothing. `--json` emits the versioned command boundary for scripts and agents: successful responses contain `contractVersion`, `ok`, and `data`; failures contain `contractVersion`, `ok`, and an `error` with a stable code, message, and any recovery guidance. Without it, output is formatted for people. `status`, `logs`, and `doctor` inspect only and do not accept `--dry-run` or `--yes`. ## Lifecycle [#lifecycle] Start, stop, and restart first inspect the selected Compose services, validate the rendered Compose configuration, make the change, then verify the desired services are stopped or healthy and running. ```bash # Inspect, then start. argus start --dry-run --json argus start --yes --json # Inspect, then stop or restart. argus stop --dry-run --json argus stop --yes --json argus restart --dry-run --json argus restart --yes --json # Non-mutating state and bounded logs. argus status argus logs --tail 200 argus logs argus --tail 200 argus logs --raw argus doctor ``` `status --json` returns `{ state: "running" | "degraded", services: { ... } }`. `state` is `running` only when every selected service is running and not unhealthy; otherwise it is `degraded`. Each service value is Docker health when present, otherwise Docker state. The human summary uses the same public state and service values. Human `logs` removes container prefixes and noisy runtime metadata; use `logs --raw` for exact bounded Docker Compose output. `logs` accepts `argus`, `postgres`, `searxng`, or `fxembed`; its tail must be a positive integer no larger than 10,000 and is time-bounded. An invalid tail is `LOG_TAIL_INVALID`; an unsupported service is `LOG_SERVICE_INVALID`. `doctor` runs independent bounded checks for Docker, the authenticated Argus API, selected storage, SearXNG, FxEmbed, and enabled source smoke targets. A check is `healthy`, `unhealthy`, or `skipped`; the complete report is healthy only when no check is unhealthy. It supplies an error code, bounded log command, and recovery text when available. Source smoke checks create an isolated temporary diagnostic watch, poll it within a deadline, then remove it; they do not use or alter normal watch records. ## Configuration and secrets [#configuration-and-secrets] Validate an intended YAML document before asking the live service to inspect, apply, and verify it. Applying an installed configuration requires the authenticated local management integration, and the service rejects a stale plan with `CONFIG_SERVICE_PLAN_STALE` or a failed verification with `CONFIG_APPLY_VERIFY_FAILED`. ```bash argus config validate /opt/argus/argus.yaml argus config show /opt/argus/argus.yaml argus config export --json argus config apply --dry-run argus config apply ``` Without a path, `config apply` uses `/opt/argus/argus.yaml`; pass an explicit path when inspecting another file. Human `config show` prints redacted YAML; `--json` returns the structured object. It structurally redacts API tokens, intelligence keys, and URL credentials in both modes. Configuration uses complete `${NAME}` references, while values live in the owner-only secrets file. Set one secret through a hidden prompt; the value is never a command argument: ```bash argus secrets set ARGUS_API_TOKEN --dry-run --json argus secrets set ARGUS_API_TOKEN --yes --json ``` The name must be an uppercase secret-like environment name or the command returns `SECRET_NAME_INVALID`. Secret values may not contain line breaks (`SECRET_VALUE_INVALID`). Argus writes `/opt/argus/secrets.env` atomically with mode `0600`, verifies the stored value, and refuses an unsafe, non-regular, non-owner-owned, or symlinked file with `SECRETS_FILE_UNSAFE`. ## Targeted repair [#targeted-repair] Repair is limited to managed `argus`, `postgres`, `searxng`, and VPS-hosted `fxembed`; any other name returns `REPAIR_SERVICE_INVALID`. Inspect the exact service before a mutation, then verify it through the doctor report: ```bash argus repair argus --dry-run --json argus repair argus --yes --json argus repair postgres --dry-run --json argus repair postgres --yes --json argus repair searxng --dry-run --json argus repair searxng --yes --json argus repair fxembed --dry-run --json argus repair fxembed --yes --json ``` PostgreSQL repair is unavailable on a SQLite deployment (`POSTGRES_NOT_SELECTED`). SearXNG repair is unavailable for disabled or external SearXNG (`SEARXNG_NOT_MANAGED`). FxEmbed repair is unavailable unless it is VPS-managed (`FXEMBED_NOT_VPS_MANAGED`). Managed SearXNG repair rewrites only its owned settings; VPS FxEmbed repair restarts only that service. A repair that cannot end with a healthy report returns `REPAIR_VERIFY_FAILED`. There is no CLI repair for external services or Cloudflare account configuration. ## Update and rollback [#update-and-rollback] An update is available only through a signed-release CLI composition. It fetches the stable release manifest and the persisted release context, verifies both, and plans the exact selected services. Review that plan before applying it: ```bash argus update --dry-run --json argus update --yes --json argus status --json argus doctor --json ``` For a non-noop SQLite plan, Argus proves the Compose-owned `argus_argus-data` volume, stops the writer, and creates a consistent snapshot under `/opt/argus/backups`. It quick-checks and hashes that snapshot and records its byte length, table counts, and exact volume identity before pulling pinned images or running migrations. It then starts Compose, checks selected service health, records the new verified state, and advances `/opt/argus/management.state` to the matching signed management CLI image. The root-owned `/usr/local/bin/argus` launcher remains immutable. An already-current release is a no-op plan but still performs health verification and can repair stale valid management state from the verified release. A missing state fails closed and requires the signed installer. The CLI does not automatically roll back a failed update; preserve the snapshot, inspect diagnostics, and choose rollback deliberately. Do not edit `management.state`: missing or malformed state fails closed before Docker runs. Rerun the signed installer to repair this launcher/state pair. The next release refuses legacy version-pinned launchers instead of upgrading them; remove an obsolete launcher explicitly before installing it. ```bash argus update --rollback --dry-run --json argus update --rollback --yes --json argus status --json argus doctor --json ``` Rollback requires the persisted verified snapshot and an exactly matching, compatible signed rollback release. For SQLite it revalidates the snapshot before stopping Argus, proves the named volume again, stages the prior database inside that volume, atomically selects it, starts the prior pinned services, verifies health, and restores `state.json`. `UPDATE_ROLLBACK_UNAVAILABLE` means required rollback support or recovery material is unavailable: this CLI integration does not support rollback, no verified rollback release is selected, the signed release context or persisted update state is missing, unreadable, or invalid, or a persisted recovery path escapes the instance root. `UPDATE_ROLLBACK_INCOMPATIBLE` means readable persisted update state has no backup, or its recorded backup/release pairing fails the required schema, manifest, or version metadata. Do not delete `/opt/argus/backups`, `update-state.json`, or `release-context.json` until recovery is complete. ## Back up and restore [#back-up-and-restore] The managed Compose SQLite database is `argus_argus-data`; confirm it with read-only Docker commands: ```bash cd /opt/argus docker compose -p argus config --volumes docker volume inspect argus_argus-data ``` Signed updates create and verify their own rollback snapshot of that volume, but retain an independently tested operator-managed Docker-volume backup for disaster recovery. Test operator recovery against a separate volume or staging instance. After any recovery, inspect the managed deployment with `argus status --json` and `argus doctor --json`. For PostgreSQL, use `pg_dump`/`pg_restore` or provider snapshots outside `/opt/argus`; after a restore, start each required role and verify its health. Argus does not create or restore PostgreSQL dumps. ## Escalation sequence [#escalation-sequence] When a command fails, retain the stable error code and start with these read-only diagnostics: ```bash argus status --json argus doctor --json argus logs --tail 200 --json ``` Use only the service-specific recovery the doctor report names. If rollback verification fails, stop mutating the host; retain `/opt/argus`, database volumes, `backups/`, and update state for investigation. # Quick start ## Outcome [#outcome] You will have a signed Argus release running on one VPS, collecting the public Argus site as a controlled Web target, and returning stored records through the authenticated API. Argus stores canonical records and revisions; it does not need an LLM for collection or querying. ## Prerequisites [#prerequisites] Use a fresh VPS running Ubuntu 22.04 or 24.04, or Debian 12 or 13, on an AMD64 or ARM64 host. You need a shell account that can use `sudo`, outbound HTTPS access, and at least 5 GiB free disk space. Docker Engine with the Docker Compose plugin may already be installed and usable by your account; otherwise the installer can install it after an explicit approval. Do not run this procedure on a workstation or an existing Argus instance. The installer writes `/usr/local/bin/argus`; onboarding owns the instance under `/opt/argus`. ## Inspect and install the signed release [#inspect-and-install-the-signed-release] First download the small installer and inspect its non-mutating report. It shows the detected platform, signed manifest URL, target wrapper path, and whether Docker is usable; it does not download or change files in inspection mode. ```bash curl -fsSLo /tmp/argus-install.sh https://argus.gpsxtre.me/install.sh ARGUS_INSTALL_INSPECT=1 sh /tmp/argus-install.sh ``` For the normal interactive path, install with the public one-command installer: ```bash curl -fsSL https://argus.gpsxtre.me/install.sh | sh ``` The installer verifies the embedded Ed25519 signature for the release manifest before it accepts the wrapper URL and SHA-256 hash, then verifies the downloaded wrapper hash before replacing `/usr/local/bin/argus`. If Docker Engine and Compose are already usable, the installer continues. If they are absent, it asks from an interactive terminal before installing them from Docker's official apt repository. A non-interactive approved install must use the downloaded script and set `ARGUS_INSTALL_DOCKER=1`; setting it to `0` instead makes a missing Docker installation fail safely. ```bash ARGUS_INSTALL_DOCKER=1 sh /tmp/argus-install.sh ``` ## Onboard one controlled Web watch [#onboard-one-controlled-web-watch] Open Argus: ```bash argus ``` Choose **Set up Argus**. The home menu is the recommended way to use Argus from a terminal; it also gives you status, readable logs, configuration, diagnostics, updates, service controls, and secrets without memorizing commands. `argus onboard` starts the same setup directly. Choose **SQLite** for this single-host start, accept port `8788`, select **Web**, and provide a watch ID such as `argus-homepage` with this controlled URL: ```text https://argus.gpsxtre.me/ ``` Leave feeds and Web search queries empty. SearXNG is asked only after at least one Web search query; choose managed SearXNG for a query or supply an external endpoint. FxEmbed is asked only when X is enabled; running it privately on the same VPS is the recommended default. Telegram accepts public channel names only. Leave OpenRouter summaries disabled unless you intend to provide an OpenRouter API key. The CLI asks for the Argus API token through a hidden prompt and keeps it out of the versioned configuration. Review the plan and confirm it. Onboarding renders the instance configuration, creates required managed services, verifies them, and writes secrets only to the owner-only instance secrets file. ## Verify [#verify] Check the deployment and diagnostics in the human-readable view: ```bash argus status argus doctor argus logs --tail 50 argus query latest-news ``` For scripts and AI agents, add `--json`. That returns the stable versioned CLI envelope; a healthy result has `ok: true`. Human logs are compact by default; `argus logs --raw` shows the exact bounded service output. The scheduler evaluates enabled watches every 30 seconds and workers poll for queued jobs every 5 seconds, so allow the configured five-minute schedule to run. To request an immediate run for the controlled watch, enter the token you created during onboarding into a shell variable, then call the authenticated endpoint from the VPS: ```bash read -rsp "Argus API token: " ARGUS_API_TOKEN; export ARGUS_API_TOKEN; echo curl -X POST -H "Authorization: Bearer $ARGUS_API_TOKEN" \ http://127.0.0.1:8788/v1/watches/argus-homepage/ingest ``` After a worker completes the queued targets, query the stored Web records. The token stays in the shell variable rather than appearing in the command. ```bash curl -H "Authorization: Bearer $ARGUS_API_TOKEN" \ "http://127.0.0.1:8788/v1/records?source=web&limit=20" unset ARGUS_API_TOKEN ``` The response contains `items` with each source URL, canonical record identity, content hash, and ingestion timestamp. An empty result means the scheduled or immediate job has not completed yet; inspect `argus doctor` before changing the host. ## Next step [#next-step] Read [core concepts](/docs/concepts) to understand the stored objects, then use [configuration](/docs/configuration) to change versioned watches and [operations](/docs/operations) for routine health and recovery work. # Security Argus protects the verified deployment path and limits the authority of its services. It does not turn a public host, operator-managed external endpoint, database backup, or Cloudflare account into a managed security boundary. ## Release trust and images [#release-trust-and-images] The installer embeds an Ed25519 public trust root. It downloads `manifest.json` and `manifest.sig`, verifies the signature and manifest shape before trusting it, then verifies the wrapper SHA-256. A malformed manifest, invalid signature, or wrapper hash mismatch stops before replacement. Onboarding and update use the verified signed manifest to select digest-pinned Argus, PostgreSQL, SearXNG, and FxEmbed images. Update also requires the persisted signed release context to match the currently deployed release. These checks establish that the CLI uses the signed release material it verified; they do not validate arbitrary images or releases an operator deploys outside the CLI. ## Secrets and credentials [#secrets-and-credentials] `/opt/argus/secrets.env` holds runtime credentials and is written atomically as an owner-owned regular file with mode `0600`. The CLI rejects unsafe permissions, symlinks, and non-files. Use `argus secrets set NAME` so the value is collected through a hidden prompt; do not put a secret in a shell command, answers file, or committed YAML. The CLI redacts known secret values from JSON and human output, redacts configured API and OpenRouter values, and removes URL userinfo in `argus config show`. Deployment errors redact registered secret values. This is output protection, not a guarantee that an operator's shell history, reverse proxy, external log collector, or third-party service never receives a secret. VPS-hosted FxEmbed needs no Cloudflare credential and exposes no host port. It still makes outbound requests to X and inherits FxEmbed/X availability and rate limits. Cloudflare-hosted FxEmbed needs an API token and account ID; use least privilege and grant access only to the intended account and Worker operations. Remove unused GHCR Docker credentials from `/opt/argus/.docker/config.json` when private image access is no longer needed. ## API exposure and authentication [#api-exposure-and-authentication] `/health` is deliberately public. When `api.token` is configured, `/v1/*` requires `Authorization: Bearer `; an absent or incorrect header receives `401`. The management configuration routes always require an exact configured Bearer token. The runtime refuses to bind a non-loopback API host without `api.token`. Loopback includes `127.0.0.1`, `::1`, `::ffff:127.0.0.1`, and `localhost`, but it is not suitable for the managed Compose topology: the management CLI and doctor reach the published container API through the host port. Docker Compose publishes the selected API port, so an Internet-facing VPS needs a required Bearer token plus an operator-controlled host firewall, reverse proxy, or private network policy. Use those network controls to limit remote access; do not rely on a loopback API host for a managed Compose instance. ## Public-source and SSRF protections [#public-source-and-ssrf-protections] Transient source primitives require a configured API token. The X primitive accepts only GET/HEAD under normalized `/2/` paths on the configured FxEmbed origin; the Web primitive exposes only JSON search on the configured SearXNG origin. Callers cannot choose an upstream host. Both paths cap redirects, request time, body size, and per-token/source rate, strip caller headers, and never persist their response. Direct Web collection accepts only HTTP(S) URLs without URL userinfo. It resolves every destination and redirect hop, rejects private, loopback, link-local, documentation, multicast, and other non-public IP ranges, and pins the connection to the approved DNS result. This prevents DNS rebinding between validation and connection. Redirects are manual, capped, cannot downgrade HTTPS to HTTP, and are revalidated at every hop. Web requests have bounded DNS, request, redirect, and body handling. The diagnostic Web target is independently validated by the same public-destination policy before Argus creates its temporary diagnostic watch. These SSRF protections apply to Argus's Web fetch path; they do not make operator-selected external SearXNG, FxEmbed, Telegram, X, OpenRouter, reverse proxies, or arbitrary infrastructure trustworthy. ## Diagnostics, backups, and operations [#diagnostics-backups-and-operations] Doctor source checks use isolated `__argus_doctor:` watches with a short expiry and cleanup path, keeping their records separate from ordinary watches. Diagnostics return stable summaries and bounded log commands rather than command stderr or external response bodies. Backups are sensitive: SQLite copies can contain all collected records, and PostgreSQL dumps can contain the same data plus credentials if an operator embeds them in a connection URI. Store backups outside the instance where appropriate, encrypt them under your backup policy, restrict access, and test recovery. Never attach `secrets.env`, tokens, signing keys, private keys, or environment dumps to support reports. ## Reporting a vulnerability [#reporting-a-vulnerability] Do not publish credentials, a working exploit, or private instance data in an issue. Send a minimal reproduction, affected version, impact, and safe contact details through the repository's private security-reporting channel if one is configured; otherwise contact the maintainers privately before disclosure. Preserve logs only after redacting secrets. # Storage schema Argus v2 defines its SQLite and PostgreSQL schemas with Drizzle. Both adapters implement the same repository contract and logical tables; SQLite stores JSON and timestamps as text while PostgreSQL uses `jsonb` and timezone-aware timestamps. ## Record graph [#record-graph] | Table | Stored data | | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `records` | One canonical source item: source/external identity, public URL, title/text/author, publication time, raw payload, metadata, content hash, and first/last seen. | | `record_watches` | Which watch and deterministic target observed a record, with first/last seen. | | `record_revisions` | Immutable content snapshots for each distinct content hash. | | `media_assets` | Ordered image, video, audio, or document URLs plus preview, MIME, dimensions, duration, alt text, and source metadata. No media bytes. | | `record_relations` | `reply_to`, `quote_of`, `repost_of`, `thread_parent`, and `links_to` edges, including unresolved external objects. | | `engagement_snapshots` | Likes, replies, reposts, quotes, views, and bookmarks at a collection time. | The record primary key is a 64-character SHA-256 identity derived from source and external ID. A source item shared by several watches therefore remains one record. Content revisions and changing engagement are separate histories. ## X conversations [#x-conversations] | Table | Stored data | | ----------------------------- | ------------------------------------------------------------------------------------------------------------ | | `conversation_tracking` | Root record, watch, state, ordering, sample/window limits, next run, stop time, burst state, and last error. | | `conversation_snapshots` | Observed and retained counts, order, pages, completeness, truncation, upstream cursor, and collection time. | | `conversation_snapshot_items` | Ranked reply-record membership and optional sort value for one snapshot. | Replies are ordinary rows in `records`; snapshot items reference them. This keeps reply provenance queryable without embedding a mutable reply array in the root post. ## Intelligence artifacts [#intelligence-artifacts] | Table | Stored data | | ------------------ | ------------------------------------------------------------------------------------- | | `artifacts` | Summary or answer content, provider, model, structured provenance, and creation time. | | `artifact_records` | Ordered canonical records used by the artifact. | | `artifact_media` | Ordered media assets considered by the model and their disposition. | Artifacts do not replace or rewrite source records. Provenance records the prompt or question, provider generation identifier, reply-sample bounds, and which media pointers were analyzed or omitted. ## Runtime state [#runtime-state] | Table | Stored data | | -------------------- | --------------------------------------------------------------------------------------------------------- | | `schema_meta` | Exact schema version. v2 refuses legacy or unversioned layouts before mutation. | | `checkpoints` | Per-target source cursor/state used by deterministic ingestion. | | `jobs` | Queued/running/completed/failed work, retry attempt, schedule, lease, and error. | | `diagnostic_watches` | Expiring isolated watches created by `argus doctor`. | | `applied_config` | Sanitized applied config snapshot, content hash, and application time. Secret values are not stored here. | Foreign keys cascade owned rows such as media, revisions, snapshots, and artifact joins. Relation targets use `SET NULL` when a referenced local record is removed so the external source identity and URL can remain as evidence. # Troubleshooting Each recovery starts with a non-mutating inspection. Run `argus` for the guided menu or use the direct commands below. Omit `--json` for readable terminal output; keep it when collecting a structured report for automation or support. Run a `--dry-run` plan before every CLI mutation that supports it, then add `--yes` only after reviewing the current plan. Stable error codes below are the CLI or doctor codes to retain when escalating. Use `argus logs --raw` only when the compact log view leaves out detail you need. ```bash # Readable first look argus status argus logs argus --tail 200 argus doctor # Readable example plan; this does not change the instance argus config apply /opt/argus/argus.yaml --dry-run ``` ### Installer failure [#installer-failure] **Inspect:** `ARGUS_INSTALL_INSPECT=1 sh /tmp/argus-install.sh` after downloading the installer with `curl -fsSLo /tmp/argus-install.sh https://argus.gpsxtre.me/install.sh`. **Meaning:** A bad Ed25519 signature, malformed manifest, or wrapper hash mismatch stops installation before the wrapper is replaced. Host prerequisite failures are reported before onboarding as `PREFLIGHT_FAILED` with individual preflight codes. **Recover:** Review the no-change inspection output, correct the reported host or release-distribution problem, then explicitly rerun `sh /tmp/argus-install.sh`. The installer has an inspection mode, not an Argus `--dry-run` flag. ### Docker unavailable [#docker-unavailable] **Inspect:** `argus doctor --json`. **Meaning:** `DOCKER_UNAVAILABLE` means the management container cannot reach Docker. During onboarding, `DOCKER_NOT_INSTALLED`, `DOCKER_COMPOSE_UNAVAILABLE`, or `DOCKER_DAEMON_UNREACHABLE` identifies the prerequisite. **Recover:** Start Docker, install the Compose plugin, or grant the operating user daemon access as the reported prerequisite requires. For a pending onboarding, keep non-secret answers in a safely permissioned file, run `argus onboard --from /path/to/answers.yaml --dry-run --json`, review its plan, then run `argus onboard --from /path/to/answers.yaml --yes --json` from a TTY to enter any required secrets through hidden prompts. For an existing instance, rerun `argus doctor --json`. ### Service is unhealthy or degraded [#service-is-unhealthy-or-degraded] **Inspect:** `argus status --json`, then `argus logs argus --tail 200 --json` and `argus doctor --json`. **Meaning:** `degraded` means at least one selected Compose service is not healthy. `ARGUS_HEALTHCHECK_FAILED`, `STORAGE_HEALTHCHECK_FAILED`, or `DOCKER_UNAVAILABLE` identifies the failed health category. **Recover:** For the managed Argus service, run `argus repair argus --dry-run --json`, then `argus repair argus --yes --json`; finish with `argus doctor --json`. For a database or SearXNG finding, use only its matching supported repair below. ### Configuration is rejected [#configuration-is-rejected] **Inspect:** `argus config validate /opt/argus/argus.yaml --json` and `argus config show /opt/argus/argus.yaml --json`. **Meaning:** Invalid YAML, schema, missing secret reference, or unsafe secret file prevents a valid configuration plan. A stale inspected plan is `CONFIG_SERVICE_PLAN_STALE`; missing live integration is `INSTALLED_CONFIG_INTEGRATION_REQUIRED`. **Recover:** Correct the non-secret YAML or set the missing secret, run `argus config apply /opt/argus/argus.yaml --dry-run --json`, then `argus config apply /opt/argus/argus.yaml --yes --json`. Argus v2 intentionally rejects `version: 1`, unversioned databases, and old table layouts before mutation. There is no compatibility migration. Point v2 at an empty SQLite file or PostgreSQL database and ingest fresh data. ### A source yields no records [#a-source-yields-no-records] **Inspect:** `argus doctor --json` and `argus logs argus --tail 200 --json`. **Meaning:** `SOURCE_DIAGNOSTIC_TARGET_NOT_CONFIGURED` means doctor has no enabled configured target to smoke-test. `SOURCE_SMOKE_FAILED` or `SOURCE_SMOKE_TIMEOUT` means the temporary isolated source check did not ingest its target; `SOURCE_SMOKE_CLEANUP_FAILED` means its temporary watch was not removed. **Recover:** Correct the enabled source, target, or schedule in YAML, then run `argus config apply /opt/argus/argus.yaml --dry-run --json` followed by `argus config apply /opt/argus/argus.yaml --yes --json`. Do not use doctor diagnostic records as ordinary watch records. ### SearXNG is unavailable [#searxng-is-unavailable] **Inspect:** `argus doctor --json` and `argus logs searxng --tail 200 --json`. **Meaning:** `SEARXNG_HEALTHCHECK_FAILED` means the endpoint did not return bounded JSON search results. `SEARXNG_NOT_MANAGED` means the selected service is external or disabled, so Argus cannot repair it. **Recover:** For managed SearXNG only, run `argus repair searxng --dry-run --json`, then `argus repair searxng --yes --json`. For external SearXNG, repair the external service under its own controls and rerun `argus doctor --json`. ### FxEmbed or X ingestion fails [#fxembed-or-x-ingestion-fails] **Inspect:** Always run `argus doctor --json` and `argus logs argus --tail 200 --json`. For VPS-hosted FxEmbed, also run `argus logs fxembed --tail 200 --json`; that service log does not exist for Cloudflare or external mode. **Meaning:** `FXEMBED_X_SMOKE_FAILED` means FxEmbed could not pass the X source smoke check. `FXEMBED_DIAGNOSTIC_SKIPPED` means X had no diagnostic target, and `FXEMBED_DISABLED` confirms X is disabled. An X `SOURCE_SMOKE_FAILED` or `SOURCE_SMOKE_TIMEOUT` means the isolated X diagnostic did not ingest its configured target. **Recover:** For VPS-hosted FxEmbed, inspect `argus repair fxembed --dry-run --json`, apply `argus repair fxembed --yes --json`, then rerun doctor. For Cloudflare or an external endpoint, repair it under that service's controls; Argus does not mutate those deployments. ### Telegram or Web parsing fails [#telegram-or-web-parsing-fails] **Inspect:** `argus doctor --json` and `argus logs argus --tail 200 --json`. **Meaning:** Telegram supports public preview channels only. A Web diagnostic can reject a private or unsafe target; public Web fetches enforce the SSRF destination policy. Source smoke failures are reported as `SOURCE_SMOKE_FAILED` or `SOURCE_SMOKE_TIMEOUT`. **Recover:** Replace the public channel or URL with a supported target, then run `argus config apply /opt/argus/argus.yaml --dry-run --json` and `argus config apply /opt/argus/argus.yaml --yes --json`. ### API returns 401 or 400 [#api-returns-401-or-400] **Inspect:** `argus config show /opt/argus/argus.yaml --json` and `argus doctor --json`. **Meaning:** `401` means a `/v1` request lacks the configured `Bearer` token or it does not match. `400` is request validation, such as an invalid `limit`, timestamp, cursor, diagnostic body, or configuration-management request; it is not repaired by restarting the service. **Recover:** For a missing or rotated API secret, run `argus secrets set ARGUS_API_TOKEN --dry-run --json`, then `argus secrets set ARGUS_API_TOKEN --yes --json`, and update the caller's `Authorization: Bearer` header. For a `400`, correct the client request and repeat it; there is no server mutation to run. ### Summary fails [#summary-fails] **Inspect:** `argus doctor --json` and `argus config show /opt/argus/argus.yaml --json`. **Meaning:** `409 {"error":"intelligence is disabled"}` means intelligence or its API key is unavailable. A malformed summary body or a limit outside 1 through 100 returns `400`. Provider failures are not converted into an Argus recovery mutation. **Recover:** Correct the intelligence configuration and secret reference, then run `argus config apply /opt/argus/argus.yaml --dry-run --json` followed by `argus config apply /opt/argus/argus.yaml --yes --json`. Correct a bad request without changing the service. ### Query fails or returns weak context [#query-fails-or-returns-weak-context] **Inspect:** `argus config show /opt/argus/argus.yaml --json`, then run `argus query latest-news --limit=50 --json`. **Meaning:** A `409` means OpenRouter intelligence is disabled or unavailable. A sourced but weak answer means the bounded recent-record selection did not contain enough matching evidence; `argus query` is not semantic retrieval or a replacement for an agent's deliberate API traversal. **Recover:** Fix intelligence configuration as for summaries. For evidence coverage, use record filters and the research skill to inspect the stored corpus and transient primitives; do not increase confidence beyond the returned data. ### X replies look incomplete [#x-replies-look-incomplete] **Inspect:** Fetch the root record and `GET /v1/records/:id/conversation` with the configured API token. **Meaning:** Reply tracking is opt-in, ends at `maxTrackingHours`, observes at most 500 distinct replies per refresh while following distinct cursors, and retains at most `maxPerPost` in the configured order. `complete` describes the bounded collection run, not all replies on X. **Recover:** Choose a longer Hot, Standard, Niche, or custom tracking horizon for future posts and apply the reviewed configuration. Use the authenticated transient X conversation primitive for read-only deeper traversal when needed; its response is not persisted. ### Update fails [#update-fails] **Inspect:** `argus status --json`, `argus doctor --json`, and `argus logs --tail 200 --json`. **Meaning:** `UPDATE_RELEASE_UNVERIFIED`, `UPDATE_STATE_UNAVAILABLE`, `UPDATE_STATE_INCOMPATIBLE`, or `UPDATE_HEALTHCHECK_FAILED` means Argus cannot establish a verified healthy candidate. `RELEASE_MANIFEST_REQUIRED` means the installed CLI lacks the signed release composition. `UPDATE_INSTANCE_NOT_ONBOARDED` means a newer stable release is available but no instance is onboarded on this host yet. **Recover:** Preserve `/opt/argus/backups`, `update-state.json`, and `release-context.json`. Run `argus update --dry-run --json`; if it returns a reviewed plan, run `argus update --yes --json`. If the candidate is unhealthy and a persisted backup is available, use the rollback sequence instead. ### Rollback is unavailable [#rollback-is-unavailable] **Inspect:** `argus update --rollback --dry-run --json`, then `argus doctor --json`. **Meaning:** `UPDATE_ROLLBACK_UNAVAILABLE` means required rollback support or recovery material is unavailable: this CLI integration does not support rollback, no verified rollback release is selected, the signed release context or persisted update state is missing, unreadable, or invalid, or a persisted recovery path escapes the instance root. `UPDATE_ROLLBACK_INCOMPATIBLE` means readable persisted update state has no backup, or its recorded backup/release pairing fails the required schema, manifest, or version metadata. Rollback dry-run fetches and selects the signed rollback release only; it does not validate the persisted backup or prove the rollback can be applied. Backup availability and compatibility are checked during apply. **Recover:** Do not delete recovery files. Run the dry-run first, then use `argus update --rollback --yes --json` only after reviewing the selected release and accepting that apply can still return an unavailable or incompatible rollback error. If it does, restore from the separately retained SQLite or PostgreSQL backup and verify with `argus status --json` and `argus doctor --json`. # Architecture The repository is organized around small packages with one direction of responsibility: source adapters create source items, the engine normalizes and classifies them, storage persists canonical records and revisions, and the app coordinates scheduled work and the API. Shared domain contracts keep those layers explicit. ## Workspace boundaries [#workspace-boundaries] | Area | Current responsibility | | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | | `packages/contracts` | Shared domain, source, storage, OCI, and management-wrapper contracts. | | `packages/config` | YAML schema, loading, reconciliation, and secret-reference sanitization. | | `packages/source-x`, `packages/source-telegram`, `packages/source-web` | Adapters and clients for the three public source families. | | `packages/engine` | Source-item ingestion, normalization, and classification. | | `packages/storage-sqlite`, `packages/storage-postgres` | Drizzle repository implementations and equivalent v2 schemas. | | `packages/scheduler` | Schedule derivation, job creation, backoff, and worker coordination. | | `packages/intelligence` | Optional OpenRouter summaries, query answers, capability discovery, and pointer-only multimodal content. | | `apps/argus` | Runtime composition, HTTP application, workers, and role selection. | | `apps/cli` and `packages/deployment` | Stable management CLI, onboarding, Compose rendering, doctor, repair, update, and external-integration boundaries. | | `packages/release` | Release manifests, signed installer and wrapper generation, and deterministic Agent Skill archives. | | `apps/web` | Next.js landing page, Fumadocs content, machine-readable routes, installer route, and Agent Skill distribution. | ## Runtime flow [#runtime-flow] Configuration is loaded through `@argus/config`; secret values are resolved at runtime rather than copied into the applied configuration. The scheduler turns watches into leased jobs. A worker selects the relevant source adapter, passes items through `@argus/engine`, and stores records, revisions, checkpoints, and artifacts through the chosen repository. The API exposes deterministic records with source links. Optional intelligence reads stored records and writes a separate summary or answer artifact. Authenticated source primitives proxy bounded FxEmbed and SearXNG reads without repository writes. SQLite runs only in the combined `all` runtime role. PostgreSQL supports the separate `api`, `scheduler`, `worker`, and `processor` roles. Keep this storage constraint intact when changing runtime composition. ## Change boundaries [#change-boundaries] Add a source by pairing an adapter package with its configuration schema, contracts, runtime registration, tests, and user documentation. Add a storage behavior to the contracts first, then implement and test it in both storage packages when it is portable. Management changes go through the CLI and deployment packages; do not make the site, skill, or prose an independent deployment implementation. The web app is a separate public surface. It derives human docs, `/llms.txt`, `/llms-full.txt`, and per-page Markdown from the same MDX source, while the installer and Agent Skill routes consume the repository release and skill packages. # Development Use Node.js 24.16.0 and pnpm 10.33.0. The root package manifest records the exact pnpm integrity-pinned version; use Corepack or another toolchain manager that honors it. ## Bootstrap [#bootstrap] ```bash corepack enable pnpm install --frozen-lockfile cp argus.example.yaml argus.yaml cp .env.example .env set -a source .env set +a pnpm argus config validate pnpm argus config apply pnpm start ``` This starts the combined local runtime. The API is at `http://localhost:8788`; use the generated local API token only in your shell or local `.env` file. ## Configuration and environment [#configuration-and-environment] `argus.yaml` is the versioned configuration input; `.env` is local-only. The supplied environment example names the runtime inputs: | Variable | Purpose | | -------------------- | ------------------------------------------------ | | `ARGUS_CONFIG` | Runtime configuration path. | | `ARGUS_API_TOKEN` | API bearer token. | | `OPENROUTER_API_KEY` | Optional OpenRouter credential for intelligence. | | `FXEMBED_ENDPOINT` | X source endpoint. | Never commit `argus.yaml`, `.env`, `secrets.env`, deployment state, test credentials, or release signing material. Do not put secret values in YAML fixtures, test transcripts, issue reports, or shell history. See [security](/docs/security) for the operator-facing secret boundary. ## Everyday commands [#everyday-commands] ```bash pnpm lint pnpm test pnpm typecheck pnpm build ``` Use `pnpm argus --help` to discover local CLI commands and `pnpm argus --version` to print the installed CLI version. For a production-like managed instance, follow the canonical [installation](/docs/install) and [operations](/docs/operations) documentation rather than manually constructing Compose files or deployment state. # Documentation The human handbook lives in `apps/web/content/docs` as version-controlled MDX. Every page needs frontmatter with a non-empty `title` and `description`. Navigation comes from the nearest `meta.json`; keep operator pages at the docs root and contributor pages under `contributing/`. Do not add aliases or copy an operator procedure into contributor or agent documentation. ## Source of truth [#source-of-truth] Write operational instructions in the existing Fumadocs pages. The web app turns that source into the rendered `/docs` handbook, `/llms.txt`, `/llms-full.txt`, and `/docs/.md`. Update the actual CLI, schema, route, or release implementation first, then document its current behavior with secret-free examples. The portable setup skill is a separate source tree at `skills/argus-setup`. Keep detailed agent request routing there, validate it with the skill scripts, and link to it from [Agents](/docs/agents) rather than duplicating its rules. ## Verify documentation changes [#verify-documentation-changes] ```bash pnpm --filter @argus/web generate pnpm vitest run apps/web/test/content.test.ts apps/web/test/routes.test.ts apps/web/test/distribution.test.ts pnpm --filter @argus/web build pnpm --filter @argus/web check:links ``` The content tests enforce route inventory, navigation separation, required frontmatter, and local Markdown links. Route and distribution tests prove the LLM, Markdown, installer, and skill endpoints. Independently, the web CI job runs the broader web suite, Lighthouse CI, and a tracked-secret-file check. Run those same checks as a contributor requirement; the repository does not configure them as a gate or ordering dependency for the Git-connected Vercel deployment. # Contributing Argus is a pnpm workspace for a self-hosted data layer and its public documentation site. Operator procedures belong in the main documentation; these pages explain how to change and verify the repository itself. ## Start here [#start-here] * [Architecture](/docs/contributing/architecture) maps the current package and application boundaries. * [Development](/docs/contributing/development) covers the pinned toolchain, local bootstrap, and safe configuration inputs. * [Testing](/docs/contributing/testing) separates ordinary, integration, live, release, and VPS verification. * [Releases](/docs/contributing/releases) describes immutable release tags, signing, smoke tests, stable-manifest promotion, and site deployment. * [Documentation](/docs/contributing/documentation) explains how the MDX, human, and machine-readable documentation surfaces stay aligned. For an installation or a running instance, use the operator-facing [quick start](/docs/quick-start), [operations](/docs/operations), and [troubleshooting](/docs/troubleshooting) guides instead. # Releases Releases are immutable signed GitHub releases. Start from a reviewed commit that passes the root quality gates, then create a SemVer tag beginning with `v`, such as `v0.1.10`. The **Signed release** workflow validates the tag and toolchain, reserves a draft release, and only publishes it after every asset is built and verified. ## Signed release contents [#signed-release-contents] The workflow builds and pushes digest-pinned Linux AMD64 and ARM64 application, management, and VPS FxEmbed images; builds the pinned Cloudflare FxEmbed worker; renders the wrapper; and creates `manifest.json` and `manifest.sig` with the Ed25519 release key. It verifies the manifest signature and shell syntax before publishing the immutable GitHub release with the installer, wrapper, manifest, public key, and FxEmbed assets. Do not rerun a tag that already has a release and do not replace its files. A failed release needs a new reviewed commit and a new version tag. ## Smoke before stable promotion [#smoke-before-stable-promotion] The **Installer smoke** workflow consumes a published signed release and tests the signed wrapper, onboarding, and `argus doctor --json` on isolated clean hosts across the supported OS, architecture, and Docker-present/absent matrix. The manual **VPS smoke** workflow verifies a selected immutable release against a controlled Web target. Only after the required smoke evidence is green may a maintainer promote the release by updating the reviewed stable bundle together: `install.sh`, `manifest.json`, and `manifest.sig` under `apps/web/public/releases/stable/`. Preserve the exact signed manifest and signature bytes, render the stable installer through the promotion tooling, and run the stable-release web tests. CI rejects a partial bundle change or an extra file under that directory; unrelated repository changes may coexist. The stable channel is intentionally mutable; each GitHub release is not. ## Public site deployment [#public-site-deployment] The public site is the Next.js project rooted at `apps/web`, with its Vercel configuration in `apps/web/vercel.json`. Its Git-connected Vercel project deploys the committed site source; do not deploy generated output or edit site assets out of band. Before merging the promotion or documentation change, run: ```bash pnpm --filter @argus/web generate pnpm vitest run apps/web/test pnpm --filter @argus/web build pnpm --filter @argus/web check:links ``` After the connected deployment completes, verify the public release link, stable manifest URL, installer, docs, and machine-readable routes from `https://argus.gpsxtre.me`. # Testing Run the smallest tier that exercises the changed boundary, then run the root quality gates before proposing a release. Tests use Vitest unless a workflow or shell smoke is named below. ## Tier 1: ordinary workspace tests [#tier-1-ordinary-workspace-tests] ```bash pnpm test pnpm typecheck pnpm lint ``` This covers package, CLI, runtime, release, and web tests. PostgreSQL repository tests are skipped unless one of these conditions is true: * `TEST_DATABASE_URL` points to an available PostgreSQL database (CI provides `postgresql://postgres:postgres@localhost:5432/postgres`); * otherwise, `ARGUS_POSTGRES_TEST=1` or `CI=true` enables the same Docker-backed Testcontainers PostgreSQL fixture, which requires Docker. The repository CI workflow supplies `TEST_DATABASE_URL`, so it uses the workflow PostgreSQL service rather than the fallback fixture. Use a disposable database only. Never commit credentialed, shared, production, or otherwise non-disposable connection strings. The workflow's isolated disposable PostgreSQL service URL is the intentional exception. ## Tier 2: web content and routes [#tier-2-web-content-and-routes] ```bash pnpm --filter @argus/web generate pnpm vitest run apps/web/test/content.test.ts apps/web/test/routes.test.ts apps/web/test/distribution.test.ts pnpm --filter @argus/web build pnpm --filter @argus/web check:links ``` Generation refreshes Fumadocs' derived source before route assertions. The web workflow also runs Lighthouse CI and rejects tracked `.env`, `state.json`, `secrets.env`, and `argus.yaml` files. ## Tier 3: controlled live integrations [#tier-3-controlled-live-integrations] The ordinary release tests validate the private FxEmbed image definition and Compose topology. The signed release workflow then builds the image for AMD64 and ARM64 from the same pinned worker bundle used by the Cloudflare artifact. The Cloudflare FxEmbed live smoke is deliberately disabled by default. Run it only against a dedicated Cloudflare test account, with all four variables supplied outside the repository: ```bash ARGUS_FXEMBED_LIVE=1 \ ARGUS_FXEMBED_LIVE_DEDICATED_ACCOUNT=1 \ ARGUS_FXEMBED_LIVE_ACCOUNT_ID=... \ ARGUS_FXEMBED_LIVE_API_TOKEN=... \ pnpm vitest run packages/deployment/test/fxembed.live.test.ts ``` It deploys and reconciles a real worker, so never point it at a shared or production account. ## Tier 4: skill and release checks [#tier-4-skill-and-release-checks] ```bash pnpm tsx scripts/skills/validate.ts skills/argus-setup pnpm tsx scripts/skills/smoke-scenarios.ts --client=fake ``` Codex and Claude smoke adapters are opt-in: each requires its CLI plus `ARGUS_SKILL_SMOKE_TEST_CREDENTIALS=enabled` and explicit disposable test credentials. The release workflow additionally runs lint, typecheck, tests, and the build before it signs and publishes release assets. ## Tier 5: clean-host smoke tests [#tier-5-clean-host-smoke-tests] The **Installer smoke** workflow runs automatically after a signed release and uses isolated Ubuntu/Debian hosts on both AMD64 and ARM64 paths. The **VPS smoke** workflow is manually dispatched with an immutable `release_tag`; its script requires `ARGUS_VPS_E2E=1`, `ARGUS_INSTALLER_URL`, `ARGUS_MANIFEST_URL`, `ARGUS_MANIFEST_ASSET_URL`, and `ARGUS_EXPECTED_VERSION`. It also accepts a controlled HTTPS URL and CI-only GitHub access variables when needed to resolve release assets. These tests create and remove disposable hosts and instance data. Run them only through their workflows or an intentionally isolated fixture environment. # Telegram v2 stores message text plus public image, video, audio, and document pointers when Telegram exposes them. It supports public announcement-channel previews only: no bot token, private chat, member-only group, or discussion replies. ## What it collects [#what-it-collects] Argus anonymously scrapes Telegram’s public channel preview for announcement posts. Each parsed post uses the canonical `https://t.me//` URL. ## Prerequisites [#prerequisites] Choose public channel usernames only. The public-preview adapter needs no Telegram account, bot token, session, or phone number. ## Configure the source [#configure-the-source] ```yaml sources: telegram: enabled: true adapter: public-web ``` `public-web` is the only supported adapter value. ## Configure watches [#configure-watches] Channel usernames contain letters, digits, and underscores, without `@` or a `t.me/` URL. ```yaml watches: - id: release-announcements schedule: "*/15 * * * *" inputs: telegram: channels: [telegram] ``` ## Validate and apply [#validate-and-apply] ```bash argus config validate /opt/argus/argus.yaml argus config apply /opt/argus/argus.yaml ``` ## Verify ingestion [#verify-ingestion] Use `argus doctor --json` and then request `GET /v1/records?source=telegram` with your API bearer token. Trigger a configured watch only when an immediate check is needed. ## Limits and safety [#limits-and-safety] Private chats, groups, direct messages, invite-only content, and content behind Telegram access controls are excluded. Preview downloads have a 20-second timeout and a 2 MiB body limit. ## Troubleshooting [#troubleshooting] * Use the channel username, not `@name`, a numeric chat ID, or the full URL. * Confirm the channel has a public preview and the source is enabled. * A Telegram HTTP failure is recorded as a failed collection job; inspect `argus logs argus`. # Web v2 extracts public page/feed image, video, audio, and document pointers in addition to text and links. Authenticated agents can call `GET /v1/primitives/web/search?q=...` for bounded transient SearXNG JSON; the response is not persisted. Agents may also pass `engines`, `categories`, `language`, `time_range`, and `pageno`; Argus rejects other parameters and always forces `format=json`. ## What it collects [#what-it-collects] Direct URLs fetch a public HTML page and extract readable text. RSS and Atom feeds yield their entries. Queries ask the trusted, managed SearXNG endpoint for discovery results; a query is not a general browser crawl. ## Prerequisites [#prerequisites] Use public `http` or `https` URLs without credentials. Configure a managed SearXNG endpoint before adding Web queries. Direct URLs and feeds do not require SearXNG. ## Configure the source [#configure-the-source] ```yaml sources: web: enabled: true searchEndpoint: http://searxng:8080 searchEndpointTrust: trusted userAgent: Argus/0.1 ``` Browser automation is not supported as a fallback. Use `searchEndpointTrust: trusted` only for an operator-controlled private service such as Argus-managed SearXNG. External endpoints default to `public`, which validates and pins every public DNS answer again on redirect hops. ## Configure watches [#configure-watches] Keep direct pages, feeds, and query discovery explicit. ```yaml watches: - id: web-signals schedule: "*/15 * * * *" inputs: web: urls: [https://example.com/news] feeds: [https://example.com/feed.xml] queries: ["critical infrastructure"] ``` ## Validate and apply [#validate-and-apply] ```bash argus config validate /opt/argus/argus.yaml argus config apply /opt/argus/argus.yaml ``` ## Verify ingestion [#verify-ingestion] Use `argus doctor --json`, then query `GET /v1/records?source=web` with bearer authentication. A `POST /v1/watches/web-signals/ingest` queues the configured targets immediately. ## Limits and safety [#limits-and-safety] Web requests enforce public-network policy: only `http`/`https`, no URL credentials, public DNS answers, and public redirects with no HTTPS downgrade. The resolver is bounded at 2 seconds; a request is bounded to 10 seconds (maximum configurable timeout 20 seconds), follows at most five redirects, and reads at most 2 MiB by default (never more than 10 MiB). SSRF targets such as loopback, private, link-local, documentation, multicast, and non-global IPv6 addresses are rejected. ## Troubleshooting [#troubleshooting] * A Web query needs `sources.web.searchEndpoint`; managed private SearXNG also needs `searchEndpointTrust: trusted`, while external public endpoints keep the `public` default. * A rejected direct URL is usually non-public DNS, a credentialed URL, a private redirect, or an HTTPS downgrade. * Check `argus logs argus` for bounded request failures or non-success feed/page responses. # X v2 stores canonical X records, public media pointers, reply/quote/repost relations, and engagement snapshots. Record identity is source-global, not tied to a watch. ## Reply tracking [#reply-tracking] Reply tracking is opt-in under `sources.x.replies`. Onboarding offers Hot (24 hours), Standard (168 hours), and Niche (720 hours). The default sample is the 50 most-liked replies Argus observed. `orderBy` also supports `newest`, `oldest`, `replies`, `reposts`, `views`, and upstream `source` order. Tracking ends at the time window, not when the sample reaches 50. Each refresh stores a conversation snapshot with observed and retained counts, ranks, completeness, and truncation. Root-post ingestion and reply refresh are independent jobs. For agent traversal, use authenticated `GET|HEAD /v1/primitives/x/2/*` calls. These bounded FxEmbed reads are transient and never ingested automatically. ## What it collects [#what-it-collects] Argus collects public X posts from configured accounts and search queries. The recommended onboarding path runs FxEmbed privately on the same VPS; Cloudflare and external endpoints are advanced alternatives. Argus does not silently use an arbitrary public FxTwitter service. ## Prerequisites [#prerequisites] Select X during `argus onboard`. The default deploys the signed FxEmbed image on the private Compose network and configures `http://fxembed:8787`. If you choose Cloudflare or external mode, supply that deployment's origin. The base URL is the origin; Argus appends `/2/...` API paths itself. ## Configure the source [#configure-the-source] ```yaml sources: x: enabled: true endpoint: http://fxembed:8787 ``` `endpoint` must be an absolute URL. Argus appends account and search paths beneath this base. ## Configure watches [#configure-watches] Accounts and queries are separate inputs: accounts track a handle’s statuses; queries send a search expression. Either list may be empty. ```yaml watches: - id: x-signals schedule: "*/10 * * * *" inputs: x: accounts: [openai] queries: ["open source AI"] ``` ## Validate and apply [#validate-and-apply] ```bash argus config validate /opt/argus/argus.yaml argus config apply /opt/argus/argus.yaml ``` ## Verify ingestion [#verify-ingestion] Run `argus doctor --json`, then query authenticated records with `GET /v1/records?source=x`. An immediate `POST /v1/watches/x-signals/ingest` queues the configured target(s). ## Limits and safety [#limits-and-safety] Only public data returned by your FxEmbed API is collected. Requests use a 20-second timeout and a 2 MiB response bound. FxEmbed response shape changes or non-success responses fail that collection job rather than silently inventing records. ## Troubleshooting [#troubleshooting] * For VPS mode, confirm the URL is exactly `http://fxembed:8787`; for an external service, use its API origin without a trailing `/2` path. * Confirm the source is enabled and the watch uses non-empty account or query values. * Inspect `argus logs fxembed` and `argus logs argus`, then run `argus doctor --json` for the configured target. ## Private account access on the local container [#private-account-access-on-the-local-container] Guest access supports public-post retrieval. Search requires an operator-provided X account; absent credentials return `X_ACCOUNT_REQUIRED`, not a successful empty search. An account configuration alone does not prove current X search/timeline compatibility. Verify both operations before enabling their watches. The local image installs the system CA bundle and retains TLS certificate verification. FxEmbed runs as UID/GID 1000 without publishing port 8787. The container reads two separate, read-only mounts: | Host path beneath the instance directory | Container input | | ------------------------------------------ | ----------------------------------------------------- | | `fxembed/credentials/credentials.enc.json` | `/run/argus-fxembed/credentials/credentials.enc.json` | | `fxembed/key/.credential-key` | `/run/argus-fxembed/key/.credential-key` | Both files must be regular files with mode 0600, readable by UID 1000. The first contains encrypted `{ciphertext, iv}`; the second is a separate 32-byte base64url key. Missing both inputs selects guest access. A partial, malformed, unreadable, or permissively readable configuration stops startup with a bounded error. At startup the entrypoint writes mode-0600 `.dev.vars` in a private temporary in-memory directory. Pinned Wrangler loads it as `ENCRYPTED_CREDENTIALS`, `CREDENTIALS_IV`, and `CREDENTIAL_KEY` bindings. Secret values never enter command arguments, Compose environment variables, YAML configuration, or image layers. Wrangler output and debug logs are suppressed; startup reports only configured, missing, or failed state. Detailed collection failures remain available as bounded Argus source/job health. No passwords or one-time codes are used. ### Encrypt and install an operator-supplied account [#encrypt-and-install-an-operator-supplied-account] Prepare a mode-0600 JSON input outside the repository with this structure. Replace placeholders privately, never in a shell command, chat, or log: ```json { "twitter": { "accounts": [{ "authToken": "VALUE_OF_auth_token_COOKIE", "csrfToken": "VALUE_OF_ct0_COOKIE", "username": "ACCOUNT_HANDLE" }] } } ``` For Buddy, the planned input is `~/.local/share/buddy-argus/credentials.json` on its VPS. That file is not provisioned by Argus; it must already contain valid operator-supplied cookies. On a Linux Docker host, from the instance directory (normally `/opt/argus`), stage it and run the image's upstream encryption tool: ```bash cd /opt/argus sudo install -d -m 0700 -o 1000 -g 1000 fxembed/credentials fxembed/key sudo install -m 0600 -o 1000 -g 1000 \ "$HOME/.local/share/buddy-argus/credentials.json" \ fxembed/credentials/credentials.json sudo docker compose run --rm --no-deps \ --volume "$PWD/fxembed/credentials:/run/argus-fxembed/credentials:rw" \ --volume "$PWD/fxembed/key:/run/argus-fxembed/key:rw" \ fxembed encrypt && sudo rm fxembed/credentials/credentials.json && sudo docker compose up -d --force-recreate fxembed ``` The `encrypt` command reuses the shipped upstream AES-256-GCM tool. It creates a private key if absent and otherwise reuses the existing key; it never prints the key or account data. Run the removal/restart steps only after encryption succeeds. The original Buddy input remains private at its original location; protect it with mode 0600 or remove it separately if you no longer need a plaintext recovery copy. The service mounts contain only encrypted credentials and the separate key after staging cleanup. Back up the two files separately and keep both outside Git and public deployment artifacts. For rotation, stage the replacement input, rerun `encrypt`, remove the staged plaintext, and recreate FxEmbed. A restart reloads bindings; no image rebuild is needed. Removing both inputs and recreating the service restores guest-only access. These instructions cover the local container, not Cloudflare credential deployment. ### Opt-in live verification [#opt-in-live-verification] Credential-free CI does not call public X. From an Argus source checkout, pipe the shipped probe into the running Argus container on its private network: ```bash docker compose -f /opt/argus/compose.yaml exec -T \ -e ARGUS_FXEMBED_LIVE=1 -e ARGUS_FXEMBED_LIVE_MODE=guest argus node --input-type=module < deploy/fxembed/check-live.mjs ``` Run before installing credentials in `guest` mode, then repeat after installation with `ARGUS_FXEMBED_LIVE_MODE=account`. Optional nonsecret probe values are `ARGUS_FXEMBED_LIVE_POST`, `ARGUS_FXEMBED_LIVE_QUERY`, and `ARGUS_FXEMBED_LIVE_HANDLE`. Account mode requires populated search and timeline results; a failure leaves authenticated access unverified. `invalid` mode expects an explicit invalid-account failure and is intended only for a separate test instance containing deliberately invalid credentials. Repeat account mode after restart and rotation. The probe prints only bounded success/failure summaries, never account values, headers, or provider response bodies.