# RebbeHub > The open, community-edited index of Chabad Torah and media: every sefer and printing, every sicha and letter, every farbrengen and recording, with scans and texts, from the Baal Shem Tov to today. Everything is readable without an account through a public API, and every change is a suggestion that people review, like a pull request (comments, reviews, #12 numbers). Problems are reported as issues (labels, assignees; public except reports of rights or of something offensive), and people have handles (@mendy) to mention and an inbox. Ids (rh-7k2m9q4d) are permanent; paths (/likkutei-sichos/12/3, /events/5742-05-10) are readable and may move. Dates are Hebrew date keys (5742-05-10 is 10 Shevat 5742; months count from Tishrei). Names are { he, en }. Words a machine read (OCR) or heard (transcription) are marked (checked: false, machine: true, origin) until a person checks them: say so when you quote them. Words whose rights forbid copies are listed but withheld. ## API - [OpenAPI 3.1](https://api.rebbehub.org/openapi.json): every route - [Interactive reference](https://rebbehub.org/developers/reference) - [MCP server](https://api.rebbehub.org/mcp): Streamable HTTP, tools search, get_item, list_children, get_text, suggest_fix, list_issues, open_issue. Reading needs no account; suggesting and opening issues need the person's account with the write scope: the server answers 401 with WWW-Authenticate so clients connect with OAuth (the person approves on rebbehub.org), or send a personal token (Authorization: Bearer rhp_…) - [Everything in one file](https://rebbehub.org/llms-full.txt) ## Docs - [Overview](https://rebbehub.org/developers): What the API, the dumps and the agent tools offer, and the promises they keep - [Getting started](https://rebbehub.org/developers/getting-started): Your first requests, in curl and TypeScript - [Tokens and signing in](https://rebbehub.org/developers/auth): Personal API tokens, their scopes, and what a token may never do - [Conventions](https://rebbehub.org/developers/api): Stability, errors, pages and cursors, caching, CORS - [The endpoints](https://rebbehub.org/developers/endpoints): A guide to the routes, by what they are for - [Sending suggestions](https://rebbehub.org/developers/suggestions): How a change is made through the API, checked and reviewed - [The data model](https://rebbehub.org/developers/data-model): Every kind of item, ids and paths, Hebrew dates, order - [Rights](https://rebbehub.org/developers/rights): What may be served and copied, and what is only listed - [Rate limits](https://rebbehub.org/developers/rate-limits): How much, per address and per token, and being a good citizen - [Webhooks](https://rebbehub.org/developers/webhooks): Every approved change posted to you, signed - [Dumps and mirrors](https://rebbehub.org/developers/dumps): The whole catalog: signed editions, the git mirror, running a mirror - [OAI-PMH and IIIF](https://rebbehub.org/developers/oai-pmh): For libraries: Dublin Core records, and IIIF manifests of scans - [AI agents](https://rebbehub.org/developers/agents): llms.txt and the MCP server: search, read and suggest - [TypeScript client](https://rebbehub.org/developers/client): @rebbehub/client, generated from the OpenAPI document ## Optional - [Download and mirror](https://rebbehub.org/mirrors): the whole catalog as signed dumps - [Source code](https://github.com/shmuky/RebbeHub) (AGPL-3.0) --- # RebbeHub for developers RebbeHub is the open, community-edited index of Chabad Torah and media: every sefer and printing, every sicha and letter, every farbrengen and recording, the scans and the texts of each. All of it is open to read through the API, with no account and no key, and everything people can do on the site - suggest a fix, fix a line, sync a paragraph - a script or an AI agent can do through the API too, under the same review. - **API**: `https://api.rebbehub.org/v1`, described in full by [OpenAPI 3.1](https://api.rebbehub.org/openapi.json) and the [interactive reference](/developers/reference). - **For AI agents**: [`/llms.txt`](/llms.txt) and an [MCP server](agents.md) at `https://api.rebbehub.org/mcp`. - **TypeScript**: [`@rebbehub/client`](client.md), a small typed client generated from the OpenAPI document. - **Everything at once**: weekly [signed dumps](../mirrors.md) (SQLite, JSON Lines, Parquet), a git mirror, [webhooks](webhooks.md) and [OAI-PMH](oai-pmh.md) for libraries. ## Where to start 1. [Getting started](getting-started.md): your first requests, in curl and TypeScript. 2. [The data model](../data-model.md): what an item is, ids and paths, Hebrew dates. 3. [Tokens and signing in](auth.md): a personal API token, to act as yourself. 4. [Sending suggestions](suggestions.md): how a change is made, checked and reviewed. 5. [Conventions](api.md): versions, errors, pages, caching, CORS; and [rate limits](rate-limits.md). 6. [Rights](../rights.md): what may be copied, and what is only listed. Read this before you copy words or files anywhere. ## The promises - **Ids never change.** `rh-7k2m9q4d` names the same item for good; paths (`/likkutei-sichos/12/3`) are readable and may move, with redirects. - **/v1 changes only by adding.** Fields and routes are added, never renamed or removed without a new version; see [Stability](api.md). - **Machine output is labelled.** Words a machine read (OCR) or heard (transcription), links it found and summaries it wrote say so (`machine: true`, `checked: false`, `origin`) until a person checks them. Keep the label wherever you show them. - **Rights are kept.** Words whose source does not allow copies are listed but withheld (`withheld` says why), everywhere: the API, the dumps, the MCP server. The code is [AGPL-3.0](https://github.com/shmuky/RebbeHub/blob/main/LICENSE); catalog facts are CC0; community text is CC BY-SA 4.0; texts from other sources keep their own licence (each text's `licence`). --- # Getting started Reading needs nothing: no account, no key. Every answer is JSON (unless a route says otherwise), and every item looks the same: ```json { "id": "rh-7k2m9q4d", "type": "event", "path": "/events/5742-05-10", "rev": 22993, "data": { "kind": "farbrengen", "date": "5742-05-10", "title": { "he": "יו״ד שבט תשמ״ב", "en": "Yud Shvat 5742" } } } ``` `type` is one of the [kinds of item](../data-model.md); `data` follows that type's JSON Schema, which `GET /v1/types` serves. ## Search ```sh curl 'https://api.rebbehub.org/v1/search?q=%D7%99%D7%95%22%D7%93+%D7%A9%D7%91%D7%98+%D7%AA%D7%A9%D7%9E%22%D7%91' ``` ```ts import { RebbeHub } from '@rebbehub/client'; const rh = new RebbeHub(); const { results, date } = await rh.search({ q: 'יו"ד שבט תשמ"ב' }); // date: { key: '5742-05-10', he: 'י׳ שבט תשמ״ב', en: '10 Shevat 5742' } ``` A query that names a Hebrew date, in Hebrew or English (`10 Shvat 5742`), also answers with the date. `GET /v1/search/moments?q=…` finds the words inside scans, texts and transcripts (a line on a page, a paragraph at the moment it is heard); `GET /v1/search/similar?q=…` searches by meaning where it is switched on. ## One item, by id or by path ```sh curl https://api.rebbehub.org/v1/entities/rh-7k2m9q4d curl 'https://api.rebbehub.org/v1/resolve?path=/events/5742-05-10' ``` ```ts const item = await rh.getItem({ id: 'rh-7k2m9q4d' }); const { id } = await rh.resolvePath({ path: '/likkutei-sichos/12/3' }); ``` Keep ids, not paths: a path may move (the old one redirects), an id never does. `?at=` gives an item as it was then; `/history` every change, who made it and what changed. ## What an item holds A sefer's sichos, a text's paragraphs, a farbrengen's recordings: ```sh curl 'https://api.rebbehub.org/v1/entities/rh-…/children?field=work&type=unit&limit=100' ``` ```ts for await (const unit of rh.all('listChildren', { id: work, field: 'work', type: 'unit' })) { console.log(unit.id, unit.data.label); } ``` Long lists come a page at a time: the answer's `next`, passed back as `cursor`, gives the next page, and is `null` on the last ([pages](api.md)). ## Events by date ```sh curl 'https://api.rebbehub.org/v1/events?within=5742-05' # Shevat 5742 curl 'https://api.rebbehub.org/v1/events?day=05-10' # 10 Shevat, every year ``` Dates are keys (`5742-05-10`: months count from Tishrei, a leap year's Adar is `06A`/`06B`); `GET /v1/dates/parse?q=…` reads a date as people write it. ## The words - A scan's text, page by page, each line with its proofread level: `GET /v1/scans//text?page=3`. A line nobody has checked is the machine's reading (`checked: false`). - A recording's transcript, each paragraph with when it is heard: `GET /v1/recordings//transcript`. - A sicha's text: its `text` items (`/v1/entities//backlinks?field=unit&type=text`) and their paragraphs (`children?field=text&type=segment`). - A file's bytes, while its rights allow: `GET /objects/`, with `X-Credit` when the rights ask for credit. ## Keeping up - `GET /v1/commits?since=`: every approved change after one, in order. - [Webhooks](webhooks.md): every approved change posted to you. - [Dumps and the git mirror](../mirrors.md): the whole catalog at once. ## Next To act as yourself - follow items, send suggestions - make a [personal API token](auth.md). Before copying words or files anywhere, read [rights](../rights.md). --- # Tokens and signing in Reading the catalog needs no account. To act as yourself - follow items, keep your places, send suggestions, fix lines, add webhooks - send a **personal API token**, or let an app such as Claude **connect with OAuth**, which you approve on the site. ## Making a token On your [account page](/account), under *API tokens*: give it a name (what will use it: "my sync script", "Claude"), pick its scopes, and optionally when it expires. The token (`rhp_` and 43 letters) is shown **once**; RebbeHub keeps only its sha256, so it cannot show it again. Lost it? Revoke it and make another. | Scope | May | | --- | --- | | `read` | read what is yours: what you follow, your places, your webhooks, report inboxes you keep | | `write` | everything you do on the site: suggestions, fixes to lines and paragraphs, sync, comments, follows, uploads, webhooks | Send it on every request: ```sh curl -H "Authorization: Bearer $REBBEHUB_TOKEN" https://api.rebbehub.org/v1/follows ``` ```ts const rh = new RebbeHub({ token: process.env.REBBEHUB_TOKEN }); ``` ## What a token is A token **is you**, no more. What it sends goes through the same checks and the same review as what you do on the site: your suggestion waits for the set's keepers, your trust level is your own, a new account's holds and daily limits on uploads apply, and a suspended account's tokens stop at once. What a token never does, whatever its scopes: sign in or change how you sign in (`/v1/auth/*`), make or revoke tokens (`/v1/tokens`), or use the stewards' tools (`/v1/admin/*`). Those happen on the site's own pages, signed in with a passkey, Google or an email link ([accounts](../accounts.md)). - A token that is unknown, revoked or expired is refused (401, `unauthorized`), never quietly read as nobody. - A `read` token that tries to change something is refused (403, `forbidden`). - Up to 20 live tokens per person. The account page shows each one's first characters, scopes, when it was made and last used. ## Connecting an app with OAuth Apps that act for you - Claude on claude.ai, Claude Code, other MCP clients - need not be handed a token. They connect with OAuth 2.1, the way the [MCP authorization spec](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization) says, and you say yes or no on rebbehub.org's own page. How to connect Claude step by step is in [AI agents](agents.md). What happens: 1. The app calls a tool that writes at `https://api.rebbehub.org/mcp` with no token (reading needs none) and gets `401` with `WWW-Authenticate: Bearer resource_metadata="https://api.rebbehub.org/.well-known/oauth-protected-resource/mcp", scope="read write"`. 2. That document ([RFC 9728](https://www.rfc-editor.org/rfc/rfc9728)) names the authorization server, `https://api.rebbehub.org`, whose metadata is at `/.well-known/oauth-authorization-server` ([RFC 8414](https://www.rfc-editor.org/rfc/rfc8414)). 3. The app uses the https address of its own description as its `client_id` (a Client ID Metadata Document, which is what Claude does by default), which RebbeHub reads, checks the redirect against, and keeps for a day; or it registers at `/oauth/register` ([RFC 7591](https://www.rfc-editor.org/rfc/rfc7591), no account needed). 4. It sends you to `/oauth/authorize` with PKCE (`S256` only) and `resource=https://api.rebbehub.org/mcp` ([RFC 8707](https://www.rfc-editor.org/rfc/rfc8707)). You land on `rebbehub.org/oauth/consent`: signed in, you see the app's name (as it calls itself), **the address it will send you back to**, and what it may do. Allow or Cancel. An address the app never registered is never redirected to. 5. The app trades the code at `/oauth/token` for an access token (`rho_…`, an hour) and a refresh token (`rhr_…`). Each refresh gives a new pair and the old refresh token stops working; a refresh token unused for 90 days expires. A code works once, within ten minutes; a code used twice ends the connection made with it. | Scope | May | | --- | --- | | `read` | always given: the catalog, and what is yours | | `write` | send suggestions, fixes, issues and comments as you; you may untick it on the consent page | A token given for `https://api.rebbehub.org/mcp` works only at the MCP server (and the calls its tools make); one given for `https://api.rebbehub.org` works on the whole API like a personal token. Like a personal token, a connection **is you**, no more, and never opens sign-in, tokens or stewards' tools, nor answers another app's request to connect. Connections are listed on your account page beside your tokens, by the app's name and where it lives; **Disconnect** ends one at once. An app may also end its own at `/oauth/revoke` ([RFC 7009](https://www.rfc-editor.org/rfc/rfc7009)). Suspending an account ends all its connections. Nothing here is ever kept in a cache: every OAuth and MCP answer says `no-store`, and none goes through Cloudflare's edge cache. What a token or a connected app sends is yours, and marked as sent by it: each suggestion, comment, issue and review keeps `via` (the token or app's name and id), and the site shows it as "Claude · for @you" with the agent mark, never as your own hands ([AI agents](agents.md)). ## Keeping it safe - Keep it out of code and out of git: an environment variable or a secret store. Anyone holding it acts as you. - Give each script or agent its own token, with `read` alone unless it must write, and an expiry when it is for a while. - Approve a connection only when you started it yourself, and check where it sends you back: `claude.ai` for claude.ai, `localhost` for Claude Code on your computer. - Revoke it on the account page the moment you think it leaked. Report a leak affecting others through [SECURITY.md](https://github.com/shmuky/RebbeHub/blob/main/SECURITY.md). ## The site's session The site itself signs in with a session cookie (HttpOnly, `SameSite=Lax`) that works only from rebbehub.org's own pages: another site's page cannot act for a signed-in reader, and the API never allows cookies across sites (CORS without credentials). Scripts and other sites use tokens. --- # Conventions What every route of `https://api.rebbehub.org` shares. The [OpenAPI document](https://api.rebbehub.org/openapi.json) describes each route; a test in the repository fails when a route is not in it. ## Stability - The API is **version 1**, under `/v1`. Within it, changes only **add**: new routes, new fields in answers, new optional parameters. Nothing is renamed or removed, and no field changes its meaning. - Write clients that ignore fields they do not know. - Should something in /v1 have to go, it is marked first: the route is `deprecated` in the OpenAPI document and answers with `Deprecation` and `Sunset` headers for at least six months, and the change is in the [changelog](https://github.com/shmuky/RebbeHub/blob/main/CHANGELOG.md). A change that cannot be made by adding becomes `/v2`, beside /v1. - `GET /v1` says the API's version; `@rebbehub/client` says the version it was generated from (`API_VERSION`). - Routes outside /v1 (`/objects`, `/dumps`, `/manifests`, `/oai`, `/mcp`) follow their own standards and keep their addresses. - The site's own sign-in (`/v1/auth/*`) and the stewards' tools (`/v1/admin/*`) are listed in the OpenAPI document for completeness, but are the site's, not for other clients, and may change with it. ## Errors Every error has one shape, with the HTTP status of its kind: ```json { "error": "not-found", "message": "rh-zzzzzzzz not found" } ``` | Status | `error` | When | | --- | --- | --- | | 400 | `bad-request` | a mistake in the request: an id that is not one, a missing field | | 401 | `unauthorized` | not signed in, or a token that is unknown, revoked or expired | | 403 | `forbidden` | not allowed: a read-only token, not a keeper of the set, a captcha not solved | | 404 | `not-found` | not there, or withheld for its rights | | 409 | `conflict`, `state` | a suggestion that clashes with main (`conflicts` lists the fields), or is not in a state for this | | 422 | `invalid` | the catalog's checks refused it (`detail` says which) | | 429 | `rate-limited` | too many requests: wait `Retry-After` seconds | | 500 | `internal` | our mistake; please report it | `message` is for people and may change; `error` is for programs and does not. ## Pages Lists that can be long come a page at a time, in a stable order: `/v1/entities`, `/v1/entities//children`, `/v1/commits` and `/v1/suggestions`. Each answer has `next`: pass it back as `cursor` for the next page; it is `null` on the last. A `Link: <…>; rel="next"` header says the same. ```sh curl 'https://api.rebbehub.org/v1/entities?type=work&limit=100' curl 'https://api.rebbehub.org/v1/entities?type=work&limit=100&cursor=c1.WyIvbGlra3V0ZWktc2ljaG9zcmgtN2sybTlxNGQiXQ' ``` Cursors are opaque: do not build or change them. They do not expire. A list keeps its order while you page through it, so a new item appears in the page it belongs to (or not, if that page is behind you). `/v1/commits` also takes `since=`, to start after any commit. Search (`/v1/search`) is ranked, not paged: ask for more with `limit`. ## Caching - Every JSON answer to a GET has an `ETag`. Send it back as `If-None-Match` and an unchanged answer is `304 Not Modified`, with no body. - Answers to anonymous requests are `Cache-Control: public, max-age=60, s-maxage=120, stale-while-revalidate=600`, and are kept at Cloudflare's edge that long (sums of the whole catalog - `/v1/stats`, `/v1/health`, `/v1/community`, `/v1/refcounts` - ten minutes; routes that never change say more: file bytes and dumps are immutable). So a change can take a couple of minutes to show to anonymous readers. The newest page of `/v1/commits` is kept only ten seconds: it and webhooks are the way to follow changes as they happen. Answers to signed-in requests are `private, no-cache` and never kept at the edge, and personal ones (`/v1/places`, `/v1/tokens`) `no-store`. `Vary: Authorization, Cookie` keeps them apart in shared caches. ## CORS Any site's pages may call the API, with a token in `Authorization`: answers carry `Access-Control-Allow-Origin: *`, and preflight requests are answered for every method. Cookies are never allowed across sites. ## Ids, dates and languages - Ids are `rh-` and Crockford base32, read forgivingly (`RH-7K2M-9Q4D`). - Hebrew date keys: `5742`, `5742-05`, `5742-05-10`, `5741-06B-14` ([data model](../data-model.md), Dates). - Names are `{ "he": "…", "en": "…" }`, Hebrew always there, English when known. ## Machine output Anything a machine made says so until a person checks it: `origin: { by, checked }` on items, `machine: true` on search moments, relations and the reviewer's advice, `checked: false` on OCR lines and transcript paragraphs. Show the label wherever you show the words. --- # The API for developers `https://api.rebbehub.org` serves the whole catalog, read without an account; `/openapi.json` lists every route. What changes the catalog needs a signed-in person: the site's session, or a personal API token made on `/account` ([auth](developers/auth.md)). The developer docs, with every route's examples and a reference to try them in, are at [rebbehub.org/developers](https://rebbehub.org/developers) ([developers/](developers/index.md)). ## Reading - `GET /v1/entities/`, `/v1/resolve?path=…`, `/v1/search?q=…`, `/v1/events?within=5742`: items, by id, path, words or date. - `GET /v1/entities//history`: every version, who made it and what changed. - `GET /v1/commits?since=`: every merge after `seq`, in order: the way to follow the catalog without webhooks. - `GET /v1/missing?kind=recordings|texts|scans`, `/v1/projects`. - `GET /v1/scans//text?page=`, `/v1/recordings//transcript`: a scan's text and a recording's transcript, machine lines marked until people check them. A scan's page gives each line's proofread level (0, 1, 2) and the scan's text layers; a transcript gives each paragraph's sync, word by word where it has word timings, and whether a person locked or checked it. - `GET /v1/scans//progress`: how far each page is proofread. - `GET /v1/recordings//hanacha`: the farbrengen's hanacha, paragraph by paragraph with where each is heard in the recording. - `GET /v1/units//printings`, `/v1/compare?a=…&b=…`: the printings of a sicha or letter whose text the catalog has (`text:`, or `scan::-` for pages of a scan), and two of them compared word by word, the Hebrew way. - `GET /v1/projects/`: a project, its progress and what is left; signed in, `POST /v1/projects//next` hands out the next item nobody holds. - Signed in: `POST /v1/scans//ocr` (your own OCR: hOCR, ALTO, or plain text with form feeds between pages), `/v1/scans//text/confirm` ("this page is right"), `/v1/recordings//sync/anchor` ("the Rebbe is saying this line now", with the moment in milliseconds) and `/v1/recordings//sync/confirm`. Each is a suggestion, reviewed. - `GET /v1/files/`: a file's size, rights and address, what was made from it (a scan's reading copy) and its page fix. - `GET /v1/page-fixes/drive/`: what a PDF on Google Drive needs to read straight - the pages to turn, each a PDF matrix (and, for a reading copy's placing, the box to cut to) - or the reading copy to open instead ([operations](operations.md)). - `/objects/`: a file's bytes, while its rights let it be served. - `/manifests//.json`: the published manifests of reading copies and page fixes. - `GET /v1/search/moments?q=…`: where the words are inside the texts: a line on a scan's page (open `/text/?page=&line=`) or a paragraph of a text or transcript, with `startMs`, when it is heard. `machine: true` until a person has checked it. - `GET /v1/search/similar?q=…&types=unit,event`: search by meaning. `available: false` until it is set up; every result is the machine's choice, and says so (`machine: true`), with its `score`. - `GET /v1/entities//relations`: an item's links both ways (cites, printed in, based on, cited by), each `machine: true` while a machine found it and no person has checked it. - `GET /v1/health`: coverage per year and set, pages nobody has checked, recordings not synced, links that do not answer, the oldest open suggestions. - `GET /v1/mirrors`, `/v1/editions`, `/v1/editions//manifest.json`, `/v1/editions//SHA256SUMS`, `/dumps//`: the git mirror, every catalog edition's dumps with their sha256, and the keys they are signed with ([mirrors](mirrors.md)). - `/manifests/iiif/.json`: a served scan as a IIIF Presentation 3 manifest (right to left, page images, the PDF as its rendering, the credit as its required statement), for any IIIF viewer. - `GET /v1/scans//pages`: a scan's pages, with their page images and thumbnails. - `GET /v1/files//similar`: other files that look like this one (the same pages, or the same recording), a machine guess. - `GET /v1/entities//linked/counts`: everything that points at an item, by type and field, with how many of each. - `GET /v1/entities//linked?field=work&type=unit&after=…&limit=…`: one of those groups in its own order (order key, date, part, page), up to 500 at a time, with `total` and `next` (null at the end). - `GET /v1/covers?ids=rh-…,rh-…` (up to 200): sefarim's covers, drawn from their title pages, while their PDFs are served or linked (a cover of a linked PDF is served; the PDF is not); `machine: true` until a person chose the page. `GET /v1/works//cover`: one sefer's cover, the page chosen, and the PDFs it may be chosen from (`linked: true` for one RebbeHub only links to). - `GET /v1/files//about`: a file's own page: rights, where it came from, what was made from it and measured in it, the covers drawn from it, and the items that use it (`usedBy.total` and the first of them). ## OAI-PMH for libraries `https://api.rebbehub.org/oai` speaks OAI-PMH 2.0 (when switched on, [deploy](deploy.md)), Dublin Core records under CC0: see [OAI-PMH and IIIF](developers/oai-pmh.md). ## Translations A unit's translation is a text of its own (`kind: translation`) with its `language`, `credit` and `licence` (none: the translator's own, CC BY-SA). `POST /v1/units//translations` with `{ language, credit, licence?, translationOf?, content, machine? }` sends one for review, a blank line between paragraphs; `machine` names the tool when a machine made it, and its paragraphs are marked so until a person checks each. `POST /v1/translations/fix` with `{ segment, content }` suggests a fix to one paragraph. Only licences that let RebbeHub keep a copy are taken (public domain, CC0, CC BY, CC BY-NC); see [rights](rights.md). ## Where you stopped `GET /v1/places`, `PUT /v1/places`, `DELETE /v1/places?kind=&key=`: where the signed-in person stopped reading (a PDF's page) and listening (a farbrengen's part and moment), the latest 60, so every device reopens there. Personal: never cached, never exported. ## Adding With a signed-in session or a token with the `write` scope: - `POST /v1/uploads/check` (`{ sha256, pageHashes?, work?, publication? }`): before an upload, whether RebbeHub has the file or one like it, and whether it looks like another scan of a printing, a new printing or a new teshura. - `POST /v1/uploads` also takes `what=hanacha` (a PDF for a farbrengen or sicha, `kind=mugah|bilti-mugah|…`) and `what=document` (`as=sefer`, `letter` or `document`, with `title`, and `set`, `author`, `genre`, `year`, `unit` as they apply); a recording or a hanacha may name a farbrengen the catalog lacks (`eventTitle`, `eventDate`) instead of `for`, and it is added with it. - `POST /v1/uploads/propose` (`{ what, name, sha256? }`): the machine's guess of where something new belongs, from the date and words in its name, and where the file already is. - `POST /v1/hanachos/text` (`{ for | eventTitle+eventDate, content, rights, language?, credit? }`): a hanacha's words, a paragraph to a segment, as a suggestion. - `POST /v1/suggestions/contents-map`: *Map pages*, what pages of a publication hold, as a suggestion. - `POST /v1/suggestions/words` (`{ entityId, change, version, segment, text, before?, kind?, language?, title?, note? }`): a page's words fixed segment by segment ([data model](data-model.md#a-pages-words)). `change` is `edit` (the segment's new `text`, as runs), `add` (a new segment after it), `remove`, or `start` (a page's first words). `before` is the segment as the person saw it: if it has changed since, the answer is 409 and nothing is overwritten. Words are only runs with the fixed marks; anything else is refused or dropped. - `POST /v1/teshuros//family-request` (no account, captcha as for reports): a family asks that a teshura stop being shown ([rights](rights.md)). ## People and conversations It works the way GitHub works. Suggestions are pull requests and Reports are issues, numbered together (`#12` is one or the other, never both; `GET /v1/threads/12` says which). Imports are not conversations and have no number. Reading needs no account (private issues aside); writing needs a signed-in session. People are named by their handle ([accounts](accounts.md#handles-and-mentions)). **People** - `GET /v1/people?q=men&thread=changeset:`: people to @mention, those already in the conversation first. - `GET /v1/people/`: a person's public page (an old handle finds them too, with `movedFrom`). - `GET /v1/threads?q=`: suggestions and issues to #mention, by number or words. **Suggestions** - `GET /v1/suggestions?state=open|closed|all&author=&reviewer=&q=`: the list, with each one's number, reviewers, approvals, comments and the issues it closes, and the open and closed counts. Without `state` the older list (`status=`) answers as before. - `GET /v1/suggestions//conversation`: the timeline in order (comments, reviews, and events: sent for review, review requested, renamed, referenced from elsewhere, merged, withdrawn, reverted), who is asked to review, and the issues it closes. - `POST /v1/suggestions//reviews`: `{ verdict: "approve" | "request_changes" | "comment", body?, comments?: [{ entity, field, body }] }`, a whole review in one: Approve merges it (as `/approve` does), Request changes sends it back (as `/send-back` does), and the comments on fields are kept with the review. - `POST /v1/suggestions//comments`: `{ body, parent?, anchor?: { entity, field } }`. - `POST /v1/suggestions//review-requests` `{ reviewers: [handle] }` (asking someone who reviewed asks again), `DELETE /v1/suggestions//review-requests/`. - `PATCH /v1/suggestions/` `{ title?, description? }`. `Fixes #12` (or `closes`, `resolves`, `סוגר`, `מתקן`…) in the description links issue 12; merging the suggestion closes it. When a suggestion is sent for review, the keepers of the sets it touches are asked to review it on their own (as CODEOWNERS are), except for a bot's suggestions. When it comes back after changes were requested, those who requested them are asked again. **Issues** - `GET /v1/issues?state=&label=&type=&set=&entity=&assignee=&author=&q=`: newest first, with the open and closed counts; `assignee=none` for those nobody has taken. - `GET /v1/issues/templates`: the kinds of issue and the words each starts with. - `POST /v1/issues` `{ title, body?, type, entityId?, labels? }`. - `GET /v1/issues/`: the issue, its timeline, the suggestions that close it, and `rights`: what the reader may do. - `PATCH /v1/issues/` `{ title?, body? }`, `POST …/state` `{ state: "open" | "completed" | "not_planned", note? }`, `PUT …/labels` `{ labels }`, `PUT …/assignees` `{ assignees }`, `POST …/visibility` `{ private }`, `POST …/comments` `{ body, parent? }`. - `GET /v1/labels`; stewards make them with `POST /v1/labels`. The suggestion list (with `state`), the issues and the inbox page like every list of the API: each answers `next`, passed back as `cursor` (`before` is still read). With a personal API token, the write scope opens issues, comments and reviews; handles are chosen on the site only (`/v1/auth/username`), never with a token. Issues are public, as on GitHub, except those about rights or offensive content, which only stewards, the set's keepers, the reporter and those assigned may read. Every report sent before issues existed stays private. `POST /v1/reports` (no account) still works and now takes a `title` too; it answers with the new issue's `number`. **Comments and the inbox** - `PATCH /v1/comments/` `{ body }` (its writer); `POST /v1/comments//resolve` `{ resolved }` (a comment on a field). - `GET /v1/inbox?filter=unread|all|mention|review_requested|assigned|…`, `GET /v1/inbox/count`, `POST /v1/inbox/read` `{ ids? | subject?: { kind, id } | all?, unread? }`. - `POST /v1/follows` takes `{ kind: "report", id }` for an issue. ## Webhooks On `/account` (*For developers: webhooks*), or `POST /v1/webhooks` with `{ "url": "https://…" }`, a signed-in person registers up to five addresses; every merge from then on is posted to each, signed with the hook's secret. The body, the signature and retries: [webhooks](developers/webhooks.md). ## Embeds Every sefer, sicha, farbrengen and set can be shown on another site (*Embed on another site*, at the foot of its page): ```html ``` A farbrengen's recordings play in place. `/embed/…` is the only page other sites may frame; every other page refuses (`frame-ancestors 'self'`). --- # Sending suggestions Nothing in the catalog changes without a **suggestion**, from the site or from the API alike. A suggestion is one or more items' new versions, with a title and a note on why. It runs the automatic checks (the item's schema, its references, its dates, duplicates), then waits for the **keepers** of the item's sets to **approve** it or **send it back** with a note; every version stays in the history, and any merge can be undone. In an *open* set, a trusted person's line fixes go live at once and are reviewed after ([governance](https://github.com/shmuky/RebbeHub/blob/main/GOVERNANCE.md)). All of this needs a [token](auth.md) with the `write` scope. ## A fix to one item Send the item's whole new `data` (read it first; change what is wrong): ```sh curl -X POST https://api.rebbehub.org/v1/suggestions/quick \ -H "Authorization: Bearer $REBBEHUB_TOKEN" -H 'Content-Type: application/json' \ -d '{ "entityId": "rh-7k2m9q4d", "title": "Fix the date", "note": "As printed in Sichos Kodesh 5742, vol. 2, p. 45", "data": { "kind": "farbrengen", "date": "5742-05-10", "title": { "he": "יו״ד שבט תשמ״ב" } } }' ``` ```ts const item = await rh.getItem({ id: 'rh-7k2m9q4d' }); const suggestion = await rh.suggestFix({ body: { entityId: item.id, data: { ...item.data, date: '5742-05-10' }, title: 'Fix the date', note: 'As printed in Sichos Kodesh 5742, vol. 2, p. 45' }, }); // { id: 812, status: 'open', checks: [...] } - waiting for review at https://rebbehub.org/review?s=812 ``` Say where it comes from in `note`: the keepers read it. ## Several items at once ```ts const draft = await rh.createSuggestion({ body: { title: 'Add the recordings of 10 Shevat 5742' } }); await rh.putSuggestionItem({ id: draft.id, body: { type: 'recording', data: { event, part: 1, title: { he: 'חלק א' }, url } } }); await rh.putSuggestionItem({ id: draft.id, body: { type: 'recording', data: { event, part: 2, title: { he: 'חלק ב' }, url: url2 } } }); const sent = await rh.submitSuggestion({ id: draft.id }); ``` `putSuggestionItem` without an `id` makes a new item (its id comes back); `data: null` deletes one. Items are checked against their type's JSON Schema (`GET /v1/types`); a failing check is in `checks`, and the keepers see it. ## Organizing the catalog Moving, renaming, ordering, merging and splitting are plans of operations, done in order, each seeing what the ones before it did. A plan becomes **one** suggestion, however many items it touches. ```ts // See where things are: the top sets, or one set or sefer, with counts. const tree = await rh.catalogTree({ root: otzros, depth: 1 }); const plan = { operations: [ { op: 'create-set', key: 'ls', name: { he: 'לקוטי שיחות', en: 'Likkutei Sichos' }, slug: 'likkutei-sichos-set', parent: sichos }, { op: 'move', items: [volume12, volume13], from: otzros, to: 'new:ls' }, { op: 'merge', from: otzrosCopy, into: likkuteiSichos }, { op: 'rename', item: likkuteiSichos, name: { en: 'Likkutei Sichos' }, slug: 'likkutei-sichos' }, ], description: 'Sorting the Otzros books into the sefarim we have', }; const preview = await rh.previewOrganize({ body: plan }); // item by item, before and after; saves nothing const made = await rh.organize({ body: plan }); // { suggestion, merged, mayApprove, preview } ``` | Operation | Does | | --- | --- | | `move { items, to, from?, mode?, position? }` | into a set (with `from`, leaving it; `mode: "only"` makes it the one set); `to: null` with `from` takes it out; a set under a set, or to the top (`to: null`); a unit to another work, a printing to a work, a scan to a printing, a recording to an event | | `move-up { items, from? }` | a set to its parent's parent; an item out of a set into that set's parent | | `rename { item, name?, slug?, path? }` | new names; a new slug or path moves the paths made from it (a sefer's units) along | | `reorder { items, parent?, position? }` | the items take their own places in the order given; with `position` (`"start"`, `"end"`, `{ after }`, `{ before }`) they go there together | | `create-set { key?, name, slug, parent?, items? }` | a new set at `/sets/`, kept like its parent; later operations name it `new:` | | `delete-set { item }` | a set that holds nothing; its path leads to its parent | | `merge { from, into }` | everything under or pointing at `from` moves to `into`, lists are joined, `from` is deleted and its paths lead to `into` | | `split { work, units? \| range, title, slug }` | some of a sefer's units into a new sefer beside it | Ordering uses fractional keys (`order`), so only the items whose place changes are changed. A set never goes under its own descendant, and an item moves only into the kind of parent it has. Every old path redirects once the suggestion is approved, and `GET /v1/entities/` of a merged item answers 404 with `detail.mergedInto`. `apply: true` approves it at once when you may approve your own (stewards); otherwise the keepers of the sets it touches review it, and stewards review changes to sets. ## Words: lines, paragraphs, sync - `POST /v1/scans//text/fix` `{ page, line, text }`: one line of a scan's text. - `POST /v1/scans//text/confirm` `{ page, fixes? }`: "this page is right" (proofread once; twice by someone else). - `POST /v1/scans//ocr`: your own OCR of a scan (hOCR, ALTO or plain text), as a layer of its own. - `POST /v1/recordings//transcript/fix` `{ segment, content }`, `/sync/anchor` `{ segment, atMs }`, `/sync/confirm`. - `POST /v1/units//translations`, `/v1/translations/fix`. ## After - `GET /v1/suggestions?author=&status=open`, and `GET /v1/suggestions/` for the review view (before and after, the checks, the reviewer's notes). - `POST /v1/suggestions//withdraw` to take it back. - Follow it (`POST /v1/follows { kind: 'changeset', id }`) to be told by email when it is decided, if you asked for email updates. - Comment on an item's talk page: `POST /v1/entities//talk { body }`. ## Being careful - Change what you know is wrong, and say how you know. A bot or an agent should say so in its token's name and in its notes. - Never send words you may not copy: see [rights](../rights.md). Machine output you send (your own OCR, a machine translation) must say so (`engine`, `machine`), and is marked until people check it. - Big imports are importers, not suggestions by hand: talk to the stewards first ([CONTRIBUTING](https://github.com/shmuky/RebbeHub/blob/main/CONTRIBUTING.md)). --- # The data model `@rebbehub/model` defines every kind of item, twice: as TypeScript types (`packages/model/src/entities.ts`) for code, and as JSON Schema (`packages/model/src/schemas/builtin.ts`) for the catalog, which checks every suggestion against them. A test keeps the two in step. ## Kinds of item **Content** - what was said or written | Type | Is | Key fields | | --- | --- | --- | | `author` | a Rebbe, a chossid, an editor, a family, an institution | `name`, `kind`, `rebbe` (1-7), `slug` | | `work` | a sefer or a series | `title`, `slug`, `authors`, `genre`, `levels` (`['volume','sicha']`), `cover` (`{ file, page }`: the title page a person chose) | | `unit` | a sicha, maamar, letter, chapter, story, diary entry | `work`, `position` (one step per level), `order`, `label`, `date`, `events` | | `event` | a farbrengen, yechidus, simcha, the writing of a letter | `kind`, `title`, `date`, `place`, `occasion` | **Print** - what exists on paper and in files | Type | Is | Key fields | | --- | --- | --- | | `publication` | a volume, kovetz, issue, booklet, teshura | `kind`, `title`, `work`, `publisher`, `date`, `printing`, `identifiers`, `simcha` | | `scan` | one PDF of one publication | `publication`, `file` (sha256), `completeness`, `preferred`, `pageLabels` | | `contents-map` | pages of a publication ↔ a unit | `publication`, `pages`, `unit` | **Text** - the words, versioned in small pieces | Type | Is | Key fields | | --- | --- | --- | | `text-layer` | the text of one scan: machine OCR, uploaded OCR, or community text | `scan`, `kind`, `engine` | | `text-page` | one page of a layer, line by line with positions | `layer`, `page`, `lines[{id,text,box}]`, `proofread` | | `text` | the text of a unit (or a recording's transcript) | `kind`, `unit`, `publication`, `language`, `licence` | | `segment` | one paragraph of a text, with a permanent id | `text`, `order`, `kind`, `content`, `proofread` | **Media** | Type | Is | Key fields | | --- | --- | --- | | `recording` | audio of an event, plus video links elsewhere | `event`, `file` or `url`, `durationMs`, `part`, `videos` | | `alignment` | a recording synced to a text | `recording`, `text`, `granularity` (word, paragraph) | | `alignment-span` | where one segment is heard | `alignment`, `segment`, `startMs`, `endMs`, `words`, `locked` | **Glue**: `relation` (based on, printed in, translation of, answer to, cites, same recording as, reproduces), `person`, `place`, `topic`, `source` (the registry of sources and their terms), `set` (what people browse; keepers and policy), and `schema` (the kinds of item themselves). Every catalog item may also carry `sets`, `externalIds` (the ids other systems know it by), `sources` (provenance), `topics` and a `note`. Anything a machine made carries `origin: { by, checked }`. ## A page's words Every item is a page people read, and many carry their own words in `body`: a letter, a chapter of a sefer, a farbrengen's outline. The words are structure, never markup (`packages/model/src/pageText.ts`, checked by the `pageText` schema): ```json { "profile": "sefaria", "versions": [ { "id": "he", "language": "he", "title": "Torat Emet", "credit": "Sefaria: Torat Emet", "licence": "cc-by-nc", "url": "https://www.sefaria.org/…", "segments": [ { "id": "14", "kind": "section", "n": 14, "text": [{ "text": "פרק יד" }], "children": [ { "id": "14.3", "kind": "verse", "n": 3, "text": [{ "text": "והנה", "marks": ["b"] }, { "text": " …" }, { "note": "n1" }, { "marker": "[כ:]" }] } ] } ], "notes": [{ "id": "n1", "kind": "note", "n": 1, "text": [{ "text": "…" }] }] }, { "id": "en", "language": "en", "segments": [ … the same ids … ] } ] } ``` - **Versions**: a page's languages and editions (a chapter's Hebrew and its English). The same place has the same segment id in every version, so versions stand side by side. - **Segments** form a tree: `section` (a title, and the segments in it), `heading`, `paragraph`, `verse` (a numbered segment), `item` (an outline's line), `note` (a footnote, in the version's `notes`). Each has an `id` that stays put, `n` when the source numbers it, `end` for a line set to the end side (a letter's date and signature), and `origin` when a machine made it. - **Words** are runs: `{ text, marks?, href? }` with only the marks `b`, `i`, `u`, `small`, `sup`, `sub`, and a link only to the web, a site path or an item id (`rh-…`, counted as a reference like any other); `{ note }` a footnote's marker; `{ marker }` a source's own marker (a page of the printed edition, a day of the study cycle); `{ br: true }`. Nothing in them is ever drawn as HTML. - **Profile**: the display rules of where the words came from. `sefaria`: sections under their titles, numbered segments (in Hebrew letters in the Hebrew), footnotes, the Hebrew and English side by side or one at a time, each version's credit and licence. `sichos-kodesh`: the texts Sichos-Kodesh publishes, paragraphs and headings, versions one at a time. `chabad-library`: chabadlibrary.org's texts, paragraphs and headings like `sichos-kodesh`, its footnotes and haoros as notes, the printed edition's old page numbers as markers, and the credit line "ספריית ליובאוויטש" ("The Lubavitch Library") linking back to the page on chabadlibrary.org, as docs/rights.md requires. `outline`: a farbrengen's contents from the Mafteiach, numbered items under titles. `plain`: what people write here. - `bodySource` records where the words were imported from, with the copy RebbeHub keeps and the rights that decide whether they are shown. Because segments are keyed lists, a change is diffed and merged segment by segment: the reviewer sees `Words › he › 14.3` before and after, and two people fixing different segments never clash. On a page's Edit tab a person clicks a segment, fixes it in place and sends it for review (`POST /v1/suggestions/words`); every segment has an anchor, so `/sefaria/…/14#s-14.3` links to one verse. Bodies were a string of wiki markup until built-in schemas version 5. `rebbehub convert-bodies` turns the catalog's over as reviewed system changes (docs/operations.md); until it has run, the API reads any it meets into structure on the way out, and a suggestion sent with one has it read on the way in. ## Ids and paths - **Ids** are permanent and opaque: `rh-7k2m9q4d` - Crockford base32, 40 bits, never reused. Read forgivingly (`RH-7K2M-9Q4D` works). Importers derive ids from their source's own ids (`idFromSeed`), so importing twice never duplicates. - **Paths** are readable and may change: `/likkutei-sichos/12/3`, `/events/5742-05-10`. An old path keeps working as a redirect. Nothing stores a path where it could store an id. ## Dates A date key is how an item is filed by Hebrew date - `5742` (a year), `5742-05` (Shevat 5742), `5742-05-10` (10 Shevat 5742), `5741-06B-14` (Purim in a leap year). Months count from Tishrei, with a leap year's Adar as `06A`/`06B` - the numbering Sichos-Kodesh's catalog uses, so keys match across both. `@rebbehub/hebrew` validates them against the calendar, sorts them, converts them to civil dates, and reads dates as people write them (`יו"ד שבט תשכ"ב`, `10 Shvat 5722`). ## Order Units and segments carry a fractional `order` key (base 62), so an item is inserted between two others without renumbering anything. ## Schema as data The built-in schemas are seeded into the catalog as `schema` items on its first commit. From then on a schema is changed like any other item - by a suggestion that a steward approves - so a new kind of item ("maaneh", "manuscript page") is a schema change, not a code release. --- # Rights Rights are RebbeHub's biggest risk, so every file carries a rights state and every export passes a gate. Code: `packages/model/src/rights.ts`, `packages/core/src/files.ts`, `packages/mirror/src/gate.ts`. ## States | State | Served? | Bytes kept in | | --- | --- | --- | | `open` | yes | the public bucket | | `credit` | yes, with its credit shown | the public bucket | | `link` | no - RebbeHub points at the source | the preservation bucket if a copy was uploaded, else nowhere | | `preserved` | no, until cleared | the preservation bucket | ## Where a file starts The stricter of what its licence allows and what was decided for its source - the same answers as Sichos-Kodesh's per-edition gate, whose four decisions map one to one (`ship`→`open`, `ship-with-credit`→`credit`, `link-only`→`link`, `local-only`→`preserved`): - public domain, CC0, facts-and-links → `open`; CC BY and CC BY-NC (Sefaria) → `credit`; - free-to-read, site terms, unknown → `link`; commercial → `preserved`; - HebrewBooks → `link` whatever else is said (a copy of a scan the jobs draw a cover from is kept in the preservation bucket, never served); - chabadlibrary.org's texts → `credit`: each page's words are kept and shown, credited to the library, with a link to its page there (a steward's decision, 2026-09-28); - the Igros app's files → `preserved` (Sichos-Kodesh decides its own apps); - hanachos and publisher scans → `link` (a copy preserved); - the old typewritten Sichos Kodesh hanachos (5710-5741) → `open`: the chozrim wrote them under no organisation, the typewritten set was printed privately in 1985, and nobody holds rights in them (the re-typed edition published since 1998 is a publisher scan); - teshuros, usually printed for free distribution → `credit`, with a fast path for families to ask for a takedown; - anything in a *locked* set → `preserved`. ## Uploads "Add a recording" (farbrengen pages) and "Add a scan" (sefer pages) take a file with a rights statement, which sets its licence: | The uploader says | Licence | Starts as | | --- | --- | --- | | I made this copy and give it freely | CC0 | `open` | | Printed or recorded for free distribution | (teshura class) | `credit` | | It is in the public domain | public domain | `open` | | I am not sure | unknown | `link`, kept privately | The bytes go to the public bucket only when the state may be served, and otherwise to the preservation bucket, which the API writes and never reads (`services/api/src/uploads.ts`). A file already known is not taken twice: the uploader is shown where it is. The file joins the catalog through a suggestion, reviewed like any other. `/add` (and "Add a hanacha" on farbrengen and sicha pages) takes a new hanacha, recording, or sefer, letter or document the same way. A hanacha's PDF whose uploader is not sure, or that was printed for free distribution, is kept privately (`link`) until a steward decides; one given freely or in the public domain is served and linked from its farbrengen's page. A hanacha's words follow the same statement: given freely or public domain, shown; otherwise kept and withheld. A sefer's **cover** is drawn from the title page of one of its PDFs, as a derivation of it: a PDF the site serves first, else one it only links to (the Otzros library on Drive, HebrewBooks). We store everything and link in public (the owner's decision, 2026-09-29): a linked PDF is fetched from its link and kept in the preservation bucket, never served; its cover is our own picture and is served, credited to the PDF's source where it has a credit, and the sefer's page links to the source for the PDF itself. A PDF in a state that keeps no copy is not fetched (the job says so). The cover is shown only while its PDF is served or linked: a takedown of the PDF (`preserved`) takes the cover down, and clearing it again brings the cover back. A new teshura ("Add a teshura" on the Teshuros set, or a scan the upload check takes for one) is always a teshura scan: `credit`, credited to the families ("משפחות כהן – לוי"), whatever rights statement comes with it. Page images and thumbnails made from a scan are derivations and follow its state. ## A family's request Every teshura page has *A family's request*: no account, works without JavaScript, captcha and rate limit as for reports. The teshura's served scans (and their page images) move to `preserved` at once, as the `system` actor in the audit log; a rights report and a `family_request` row are kept for the stewards, who may restore the state with `setRights` if the request was not from the family. Nothing is deleted (`familyRequest` in `packages/core/src/print.ts`). Requests about anything other than a teshura use the general report form. ## Translations "Add a translation" on a unit's page takes the words themselves, so it takes only what may be copied, and says whose it is: | The translator says | Licence | Served | | --- | --- | --- | | Mine, I translated it | none (community text, CC BY-SA) | yes, credited to them | | Public domain | public domain | yes | | CC0 / CC BY / CC BY-NC | as given | yes, with the credit given (Sefaria's are CC BY-NC) | A publisher's all-rights-reserved translation (Kehot's, a site's terms) is never pasted in: it is listed as a copy elsewhere, a link. A machine translation says which tool made it and is marked as machine text, paragraph by paragraph, until a person checks each one. ## Changing it Only stewards change a file's state (`setRights`), and every change is in the audit log. A takedown moves a file to `preserved`: it stops being served at once and is never deleted. Anyone asks for one at `/takedown`, with no account; a steward answers within two weeks and takes each file the request points at down in one click from `/admin` ([accounts](accounts.md#stewards-and-platform-admins)). ## Exports The git mirror and the dumps carry catalog facts always, file hashes but never files, and words only when their rights allow: a text copied from a source keeps that source's licence (site-terms and commercial texts are listed as withheld), community text is CC BY-SA, and OCR pages follow their scan's file. The Parquet dump carries exactly what the SQLite and JSON Lines dumps do. Machines that read the words (search by meaning, citations) send words to Workers AI only when they may be exported; citations found are facts (a link and the reference as written), so they are proposed from any text. --- # Rate limits The API is free and open; the limits keep it so for everyone. | Who | Allowed | | --- | --- | | Each address (no token) | 300 requests a minute | | Each API token | 1,200 requests a minute, wherever it is sent from | | Searching, each address (no token) | 60 searches a minute (`/v1/search`, `/v1/search/moments`, `/v1/search/similar`), counted on top of the above | | Google Drive files, each address | 120 a minute (`/v1/drive/`, the site's reader and player), counted on top of the above | - Every answer says the policy: `RateLimit-Policy: "address";q=300;w=60, "token";q=1200;w=60`. - Past it, the answer is `429` with `{ "error": "rate-limited" }` and `Retry-After: 60`. Wait that long; do not retry at once. - Reads anyone may make are answered from Cloudflare's edge for a minute or two (longer for the sums of the whole catalog); an answer from there does not count against you, and says `Cf-Cache-Status: HIT`. - File bytes (`/objects/`, asked for in many small ranges while a recording plays) are not counted. - A request with a bad token counts against its address, so guessing tokens is slow. - Some things have limits of their own, whatever the rate: reports and takedown requests (10 an hour from an address without an account), sign-in links by email, and uploads (a new account waits a day; so many files and bytes a day, [accounts](../accounts.md)). ## Being a good citizen - Send a token when you run a script: you get your own allowance, and a way for us to reach you if something goes wrong. - Use `If-None-Match` ([caching](api.md)): a 304 costs almost nothing. - For the whole catalog, take a [dump](../mirrors.md) instead of walking every item; to keep up after it, follow `/v1/commits` or a [webhook](webhooks.md). - Say who you are in `User-Agent` (`my-app/1.0 (+https://…)`). ## Running your own On Cloudflare Workers the limits are Cloudflare's rate limiting bindings (`RATE_LIMIT_ADDRESS`, `RATE_LIMIT_TOKEN`, `RATE_LIMIT_SEARCH`, `RATE_LIMIT_DRIVE` in `services/api/wrangler.toml`, [configuration](../configuration.md)); without them nothing is counted. `npm run dev:api` counts nothing unless `RATE_LIMIT=1`. --- # Webhooks Every change approved into the catalog can be posted to an address of yours, signed, within a few minutes. ## Adding one On your [account page](/account) (*For developers: webhooks*), or with a [token](auth.md) that may write: ```sh curl -X POST https://api.rebbehub.org/v1/webhooks -H "Authorization: Bearer $REBBEHUB_TOKEN" \ -H 'Content-Type: application/json' -d '{ "url": "https://example.org/rebbehub-hook" }' ``` ```json { "id": 3, "url": "https://example.org/rebbehub-hook", "secret": "…" } ``` The `secret` is shown this once: keep it to check signatures. Up to five addresses per person; `GET /v1/webhooks` lists yours, `DELETE /v1/webhooks/` removes one. ## What arrives Every merge from then on, in order and at least once, every few minutes: ```http POST Content-Type: application/json X-RebbeHub-Signature: sha256= { "commits": [ { "seq": 9, "at": "…", "message": "…", "author": "…", "mergedBy": "…", "changes": [ { "id": "rh-…", "type": "event", "path": "/events/…", "rev": 22993, "data": { … } } ] } ] } ``` - `data: null` means the item was deleted. - Words withheld for rights are left out, as in `/v1/commits`. - Answer with a 2xx status. Anything else is tried again; after 20 failures in a row the hook is switched off (the account page says so). - Keep the last `seq` you handled: a commit may come twice. ## Checking the signature ```ts import { createHmac, timingSafeEqual } from 'node:crypto'; function signed(body: string, header: string | null, secret: string): boolean { const expected = `sha256=${createHmac('sha256', secret).update(body).digest('hex')}`; return header !== null && header.length === expected.length && timingSafeEqual(Buffer.from(header), Buffer.from(expected)); } ``` Check it against the raw body, before parsing it. ## Without a server `GET /v1/commits?since=` answers the same commits, a page at a time ([pages](api.md)): poll it instead when you cannot receive posts. --- # Mirrors: keeping a copy of RebbeHub The plan's third promise is that the catalog belongs to the community and can never be locked away. Everything needed to keep a full copy is public, at addresses that do not change, and anyone may serve that copy from a site of their own. The site's [Download and mirror](https://rebbehub.org/mirrors) page shows it all in words. ## What there is | What | Where | Licence | | --- | --- | --- | | The git mirror: one JSON file per item, texts as Markdown, sync as WebVTT, one git commit per approved change | `git` in `/v1/mirrors` | facts CC0, community text CC BY-SA, source texts their own | | Catalog editions: the whole catalog at one commit, weekly, as SQLite and JSON Lines, plus the Sichos-Kodesh release | `/v1/editions`, files at `/dumps//` | the same | | Each edition's signed manifest and checksums | `/v1/editions//manifest.json`, `/v1/editions//SHA256SUMS` | CC0 | | The release keys (public halves) | `keys` in `/v1/mirrors` | - | Files themselves (scans, recordings) are not in the dumps or the git mirror, only their hashes and rights; words whose source forbids copies are left out and listed as withheld ([rights](rights.md), Exports). `GET https://api.rebbehub.org/v1/mirrors` answers with all of it: ```json { "git": ["https://github.com/rebbehub/catalog.git"], "dumps": "https://api.rebbehub.org/dumps", "keys": [{ "alg": "ed25519", "keyId": "…", "publicKey": "…" }], "editions": [{ "tag": "2026.40", "commit_seq": 1234, "created_at": "…", "dumps": { "files": [{ "name": "rebbehub-2026.40.sqlite", "bytes": 123, "sha256": "…", "url": "https://api.rebbehub.org/dumps/2026.40/rebbehub-2026.40.sqlite" }], "manifest": "https://api.rebbehub.org/v1/editions/2026.40/manifest.json", "sha256sums": "https://api.rebbehub.org/v1/editions/2026.40/SHA256SUMS", "signature": { "alg": "ed25519", "keyId": "…" } } }] } ``` An edition is tagged before its dumps are made; until they are, its `dumps` is `null`. ## Running a mirror ```sh git clone https://github.com/shmuky/RebbeHub && cd RebbeHub && npm ci npm run rebbehub -- mirror-pull --out /srv/rebbehub --key ``` `mirror-pull` reads `/v1/mirrors`, and for every edition (or `--tag 2026.40`, or `--tag latest`) fetches the signed manifest, checks its Ed25519 signature against the key you pinned, downloads each file, checks its sha256 against the manifest, and only then moves it into place. A file that does not match is never kept, and the run fails loudly. Editions already there and whole are left alone, so run it from cron as often as you like: ```cron 17 * * * * cd /opt/RebbeHub && npm run rebbehub -- mirror-pull --out /srv/rebbehub --key >> /var/log/rebbehub-mirror.log 2>&1 ``` The folder it leaves is ready to serve as it is, with any web server: ``` /srv/rebbehub/editions.json what is here, newest first, and where the git mirror is /srv/rebbehub/2026.40/manifest.json the edition's signed manifest /srv/rebbehub/2026.40/SHA256SUMS for sha256sum -c /srv/rebbehub/2026.40/rebbehub-2026.40.sqlite … ``` Pin the key. Without `--key`, `mirror-pull` trusts the keys the API names and says so; that checks the files arrived whole, but not that they came from RebbeHub's release key. `--api` pulls from another mirror's API instead; `--allow-unsigned` keeps editions whose manifest is unsigned (never by default). The git mirror is mirrored like any git repository: ```sh git clone --mirror https://github.com/rebbehub/catalog.git && cd catalog.git && git remote update # from cron ``` By hand, without this repository, one edition: ```sh curl -O https://api.rebbehub.org/v1/editions/2026.40/SHA256SUMS curl -O https://api.rebbehub.org/dumps/2026.40/rebbehub-2026.40.sqlite # and the other files it lists sha256sum -c SHA256SUMS ``` ## Being listed A mirror that keeps a full copy and wants to be listed on `/mirrors` asks in an issue; a steward adds it to the API's `others` list. ## Publishing (stewards) Each week ([operations](operations.md), Editions and dumps): ```sh rebbehub edition --by shmuly rebbehub dump --tag 2026.40 --out dumps/2026.40 --key release-key.json --upload ``` `--upload` puts the files and the manifest in the public bucket at `dumps//` (with `CLOUDFLARE_ACCOUNT_ID` and `CLOUDFLARE_API_TOKEN`) before recording the manifest, so an edition lists its files only once they can be fetched. The API serves only the files an edition's manifest names, and they never change. The API's settings for mirrors are public values, in `services/api/wrangler.toml` `[vars]` once they exist: | Variable | Is | | --- | --- | | `CATALOG_GIT_URL` | where the git mirror is cloned from (comma separated for several) | | `RELEASE_PUBLIC_KEYS` | the release keys' public halves, base64 (comma separated); printed by `rebbehub keygen` | | `DUMPS_BASE_URL` | where dumps are served, when not the API's own `/dumps` | Until they are set, `/v1/mirrors` lists no git mirror and no keys, and the page says they are coming. The secret half of the release key never goes in any file in this repository or in Cloudflare; it stays with the steward who signs. --- # OAI-PMH for libraries `https://api.rebbehub.org/oai` speaks [OAI-PMH 2.0](https://www.openarchives.org/OAI/openarchivesprotocol.html), so library catalogs and aggregators can harvest RebbeHub like any repository (it is on once the server names its administrators' address, `OAI_ADMIN_EMAIL`; see [configuration](../configuration.md)). - **Records**: sefarim, sichos and letters, farbrengens, printings and recordings, as Dublin Core (`oai_dc`). Records are CC0. - **Identifiers**: `oai:rebbehub.org:rh-…`, the item's permanent id. - **Dates**: a record's datestamp is when the commit that last changed it was made; harvest by `from` and `until` (day granularity or seconds). - **Sets**: by kind (`type:unit`, `type:event`, `type:publication`…) and by RebbeHub set (`set:rh-…`); `ListSets` lists them. - **Deletions** are kept (`deletedRecord: persistent`): a deleted item answers with a deleted header. - **Pages**: a hundred records at a time, with a `resumptionToken`. - Words withheld for rights are never in a record. ``` /oai?verb=Identify /oai?verb=ListSets /oai?verb=ListRecords&metadataPrefix=oai_dc&from=2026-09-01 /oai?verb=ListRecords&metadataPrefix=oai_dc&set=type:unit /oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:rebbehub.org:rh-7k2m9q4d ``` `POST /oai` takes the same arguments as a form. ## IIIF Every scan RebbeHub serves has a IIIF Presentation 3 manifest, `https://api.rebbehub.org/manifests/iiif/.json`: right to left, its page images, the PDF as its rendering, the credit as its required statement. Open it in any IIIF viewer (Mirador, Universal Viewer). --- # AI agents RebbeHub is ready for AI agents two ways: plain text to read ([`/llms.txt`](/llms.txt), [`/llms-full.txt`](/llms-full.txt)), and an MCP server with tools to search, read and suggest. ## llms.txt `https://rebbehub.org/llms.txt` says in a page what RebbeHub is and where everything is ([llmstxt.org](https://llmstxt.org)); `/llms-full.txt` is every developer page and the list of API routes in one file, to give an agent as context. The API has its own short `https://api.rebbehub.org/llms.txt`. ## The MCP server The [Model Context Protocol](https://modelcontextprotocol.io) server speaks Streamable HTTP (versions 2024-11-05 to 2025-11-25): POST a JSON-RPC message, get JSON back. It keeps no sessions and opens no event stream. It has one address, `https://api.rebbehub.org/mcp`: - **Reading needs no account.** `search`, `get_item`, `list_children`, `get_text` and `list_issues` answer anyone. - **Writing asks for you when it is needed.** A tool that writes (`suggest_fix`, `open_issue`) called without sign-in answers HTTP `401` with `WWW-Authenticate: Bearer resource_metadata="https://api.rebbehub.org/.well-known/oauth-protected-resource/mcp", scope="read write"`, the MCP authorization spec's step-up, so the client asks you to sign in and tries again. Connected for reading only, it answers `403` with `error="insufficient_scope"` and the scope it needs. A token that has ended or was revoked answers `401` with `error="invalid_token"`. - **Sign-in is OAuth 2.1** (you approve the app on rebbehub.org), or a personal token sent as `Authorization: Bearer rhp_…`. | Tool | Does | | --- | --- | | `search` | items by name or Hebrew date (`where: "names"`), or the lines and paragraphs that hold the words (`where: "words"`) | | `get_item` | one item by id or path, with all its data | | `list_children` | what an item holds, in order: a sefer's sichos, a text's paragraphs, a farbrengen's recordings, a set's items; a page at a time | | `get_text` | the words of a sicha, a scan's page or a recording's transcript; machine words marked `[machine]` | | `suggest_fix` | a correction to one item, as a suggestion for review, under your name (`write`) | | `list_issues` | issues people opened, open ones first, or about one item | | `open_issue` | report a problem for people to look into, under your name (`write`) | | `get_tree` | the catalog as a tree: the top sets, or one set or sefer, with the sets and items under it and how much each holds | | `preview_organize` | what a plan of organizing operations would change, item by item, the paths that redirect; saves nothing | | `organize` | a whole plan of organizing operations as one suggestion (needs `write`) | | `move_items` | sefarim into or out of sets, a set under another set or to the top, sichos to another sefer, a printing, scan or recording to another parent (needs `write`) | | `move_up` | a set to its parent's parent; a sefer out of a set into that set's parent (needs `write`) | | `rename_item` | a new name in Hebrew and English, and optionally a new slug or path; old paths redirect (needs `write`) | | `reorder_children` | put a sefer's sichos, a set's sefarim or sets in a new order (needs `write`) | | `create_set` | a new set, under another or at the top, with items moved in at once (needs `write`) | | `delete_set` | remove a set that holds nothing (needs `write`) | | `merge_items` | merge a duplicate into the item kept: its children and links move over, its paths redirect (needs `write`) | Every tool calls the API itself, as you, so an agent reads exactly what anyone reads: words withheld for rights stay withheld, and a fix it suggests is a Suggestion reviewed like anyone's. Nothing in the catalog changes until a keeper approves it. ### Organizing the catalog The organizing tools (`move_items` to `merge_items`, and `organize` for several steps at once) each make **one** suggestion, however many items it touches: a sefer renamed with its three thousand sichos' paths is one suggestion. Nothing changes until a keeper of the sets it touches approves it (stewards approve changes to sets themselves). Every old path redirects once it is approved; a merged item's paths lead to the item it was merged into. Look first with `get_tree`, and send the plan to `preview_organize` to see the change before making it. The same plans go to the API as `POST /v1/organize` ([suggestions](suggestions.md#organizing-the-catalog)). ### What an agent sends shows as the agent's A suggestion, comment, issue or review sent with a token or a connected app is kept with what sent it (`via`: the token or app's name and id). The site shows it everywhere the author shows as the agent's work for you, with the agent mark: **Claude · for @you** (**Claude · בשביל @you**), on the suggestion, the lists, the item's history and your page, and a suggestion says an agent sent it until a person reviews it. The API returns `via` on suggestions, their conversation, issues, history and commits; it is `null` for what a person did on the site. ### Connecting from claude.ai (and the Claude apps) The recommended settings in claude.ai's connector form: | Field | Choose | | --- | --- | | Name | RebbeHub | | Remote MCP server URL | `https://api.rebbehub.org/mcp` | | Authentication | **Sign in when needed** (reading works at once; you sign in the first time Claude suggests a fix). **Sign in now** works too. **No sign-in** reads only. | | OAuth client | **Use Claude's published identity (CIMD)**, recommended. **Register automatically (DCR)** works as a fallback. No client id or secret of your own is needed. | 1. On claude.ai, open **Settings → Connectors** and choose **Add custom connector**, with the settings above. 2. Choose **Add** (and **Connect**, with *Sign in now*; with *Sign in when needed*, Claude asks the first time a tool writes). 3. When asked to sign in, Claude opens rebbehub.org: sign in (with your passkey, Google or an email link) if you are not already, check that the app is Claude and that it sends you back to `claude.ai`, and choose **Allow**. Untick *Send suggestions as you* to let it only read. 4. In a chat, turn RebbeHub on from the tools menu. Ask it to find something, then to suggest a fix; the suggestion appears at `rebbehub.org/review` as Claude's, for you, for a keeper to approve. The connection shows on your [account page](/account) under *Developers*, with Claude's name; **Disconnect** there ends it at once. Claude Desktop and the mobile apps use the same connectors. ### Connecting from Claude Code ```sh claude mcp add --transport http rebbehub https://api.rebbehub.org/mcp ``` Reading works at once. To write, run `/mcp` in Claude Code, pick rebbehub and **Authenticate** (or let it ask when a tool writes): your browser opens rebbehub.org to approve it, as above (it comes back to `localhost`, which the page points out). Claude Code keeps and refreshes the tokens itself. With a personal token instead (a script, a server, a client without OAuth): make one with `write` on the account page, named for the agent, and send it as a header: ```sh claude mcp add --transport http rebbehub https://api.rebbehub.org/mcp --header "Authorization: Bearer $REBBEHUB_TOKEN" ``` ### Other clients Clients that speak MCP's authorization (Cursor, VS Code, the MCP Inspector, the SDKs) find everything from the `401` a writing tool answers, or from `/.well-known/oauth-protected-resource/mcp`: its `WWW-Authenticate` names the [resource metadata](auth.md#connecting-an-app-with-oauth), which names the authorization server. Clients that take a JSON config and no OAuth send a token: ```json { "mcpServers": { "rebbehub": { "type": "http", "url": "https://api.rebbehub.org/mcp", "headers": { "Authorization": "Bearer rhp_…" } } } } ``` ### By hand ```sh curl -s https://api.rebbehub.org/mcp -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"search","arguments":{"query":"יו\"ד שבט תשי\"א"}}}' ``` ## Rules for agents - Quote the Rebbe's words only from checked text. A line marked `[machine]` (or `checked: false`, `origin` without `checked`) is a machine's reading or hearing; say so if you use it. - Link to what you cite: every result carries its page's address. - Suggest fixes only with a source, in the note; people review every one. - Organize in small, clear steps with a note saying why; preview a merge before sending it, and merge only what is truly the same item. - Keep to [rights](../rights.md): do not copy withheld words from elsewhere into a suggestion. --- # The TypeScript client `@rebbehub/client` (in this repository at `packages/client`) is a small typed client for the API, generated from its OpenAPI document: one method per operation, named by its `operationId`, taking one object of its path and query parameters (and `body`), and answering the typed JSON. It has no dependencies and runs wherever `fetch` does: Node 18+, browsers, Workers, Deno, Bun. Its licence is the site's own, [AGPL-3.0](https://github.com/shmuky/RebbeHub/blob/main/LICENSE), and the package carries a copy. ```ts import { RebbeHub, RebbeHubError } from '@rebbehub/client'; const rh = new RebbeHub({ token: process.env.REBBEHUB_TOKEN, // optional: reading needs none userAgent: 'my-app/1.0', }); const { results } = await rh.search({ q: 'לקוטי שיחות חלק יב', limit: 5 }); const item = await rh.getItem({ id: results[0]!.id }); const page = await rh.scanText({ id: 'rh-…', page: 3 }); ``` ## Pages `pages` walks a paged list page by page; `all` gives its items one by one, fetching pages as they are needed: ```ts for await (const commit of rh.all('listCommits', { since: 1200 })) { console.log(commit.seq, commit.message); } ``` ## Errors Every error the API answers is thrown as a `RebbeHubError`, with the HTTP `status`, the error's `code` (`not-found`, `rate-limited`…), `message`, `detail`, and for a 429 `retryAfter` in seconds: ```ts try { await rh.getItem({ id: 'rh-zzzzzzzz' }); } catch (error) { if (error instanceof RebbeHubError && error.code === 'not-found') … } ``` ## Options | Option | Default | | --- | --- | | `baseUrl` | `https://api.rebbehub.org` (`http://127.0.0.1:8787` for `npm run dev:api`) | | `token` | none | | `fetch` | the global `fetch` (pass a Worker's service binding, or a test's) | | `userAgent` | added to `User-Agent` | ## Keeping it in step `packages/client/src/generated.ts` is written by `npm run generate -w @rebbehub/client` from `services/api/src/openapi.ts`, and a test fails when it is not up to date, so the client changes with the API in the same pull request. The site's own sign-in and the stewards' tools are left out: a token cannot use them. --- # Every route (RebbeHub API 1.0.0, https://api.rebbehub.org) - `GET /` root: Redirects to /v1. No account needed. - `GET /v1` about: About this API: its version, the latest commit, where the docs are. No account needed. - `GET /openapi.json` openapi: This document. No account needed. - `GET /v1/stats` stats: How many items of each type, and the latest commit. No account needed. - `GET /v1/community` community: The community page in numbers: the latest merges, reports and suggestions waiting, people, gaps. No account needed. Parameters: limit? (query). - `GET /v1/health` health: The health of the catalog: coverage per year and set, unchecked pages, unsynced recordings, dead links, the oldest open suggestions. No account needed. Parameters: limit? (query). - `GET /v1/types` types: Every kind of item and its JSON Schema. No account needed. - `GET /v1/entities` listItems: Items on main, by type and set, in path order, a page at a time. No account needed. Parameters: type? (query), set? (query), after? (query), limit? (query), cursor? (query). Paged (cursor). - `GET /v1/entities/batch` getItems: Several items at once, in the order asked (missing ones left out). No account needed. Parameters: ids (query). - `GET /v1/entities/{id}` getItem: One item, on main or as of a commit. No account needed. Parameters: id (path), at? (query). - `GET /v1/entities/{id}/children` listChildren: An item's children in their own order (a work's units, a text's paragraphs), a page at a time. No account needed. Parameters: id (path), field (query), type (query), after? (query), limit? (query), cursor? (query). Paged (cursor). - `GET /v1/entities/{id}/linked/counts` linkedCounts: What points at an item, by type and field, with how many of each. No account needed. Parameters: id (path). - `GET /v1/entities/{id}/linked` listLinked: One group of what points at an item, in its own order, a page at a time, with the total. No account needed. Parameters: id (path), field (query), type? (query), after? (query), limit? (query), cursor? (query). Paged (cursor). - `GET /v1/entities/{id}/history` itemHistory: Every merged change to an item, newest first: who, when, and what changed field by field. No account needed. Parameters: id (path). - `GET /v1/entities/{id}/backlinks` itemBacklinks: Items that point at this one. No account needed. Parameters: id (path), field? (query), type? (query). - `GET /v1/entities/{id}/relations` itemRelations: An item's links both ways: cites, printed in, based on, cited by. No account needed. Parameters: id (path). - `GET /v1/revisions/{rev}` getRevision: One stored version of an item. No account needed. Parameters: rev (path). - `GET /v1/resolve` resolvePath: The item at a readable path (an old path answers with where it moved). No account needed. Parameters: path (query). - `GET /v1/events` listEvents: Farbrengens and other events by Hebrew date, each with how many recordings it has. No account needed. Parameters: within? (query), day? (query), dates? (query), missing? (query), limit? (query). - `GET /v1/refcounts` refCounts: How many items point at each item through a field (field=work&type=unit: each work's units). No account needed. Parameters: field (query), type? (query). - `GET /v1/sitemap` sitemaps: Every sitemap there is: each kind of item with a page of its own, in pages of pageSize items (id order), with when each page last changed. No account needed. - `GET /v1/sitemap/{type}/{page}` sitemapPage: One sitemap's items: their ids, paths and when each last changed. No account needed. Parameters: type (path), page (path). - `GET /v1/works/{id}/outline` workOutline: A work's volumes (its top-level parts), with how many units each holds. No account needed. Parameters: id (path). - `GET /v1/works/{id}/parts/{part}` workPart: The units of one volume of a work. No account needed. Parameters: id (path), part (path), limit? (query). - `GET /v1/commits` listCommits: Every merge to main after a given one, in order, with what each changed: the way to follow the catalog without webhooks. No account needed. Parameters: since? (query), limit? (query), cursor? (query). Paged (cursor). - `GET /v1/dates/parse` parseDate: Read a Hebrew date as people write it. No account needed. Parameters: q (query). - `GET /v1/search` search: Search names, text and dates, in Hebrew or English. No account needed. Parameters: q (query), type? (query), limit? (query). - `GET /v1/search/moments` searchMoments: Where the words are: lines on scans' pages (open at the line) and paragraphs of texts and transcripts (open at the moment heard). No account needed. Parameters: q (query), limit? (query). - `GET /v1/search/similar` searchSimilar: Search by meaning (embeddings); every result is the machine's guess. No account needed. Parameters: q (query), types? (query), limit? (query). - `GET /v1/scans/{id}/text` scanText: A page of a scan's text: the community page, else the seed layer's; each line with its proofread level. No account needed. Parameters: id (path), page? (query). - `GET /v1/scans/{id}/progress` scanProgress: How far each page of a scan is proofread (0, 1 or 2). No account needed. Parameters: id (path). - `POST /v1/scans/{id}/text/fix` fixScanLine: Fix one line of a scan's text (a suggestion). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/scans/{id}/text/confirm` confirmScanPage: This page is right: raise it a proofreading level, with any fixes (a suggestion). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/scans/{id}/ocr` uploadOcr: Upload your own OCR of a scan (hOCR, ALTO, or plain text with form feeds between pages) as a new layer. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/scans/{id}/text/seed` seedScanText: Keepers: seed the community text from this OCR layer (checked lines are kept). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `GET /v1/units/{id}/printings` unitPrintings: The printings of a unit whose text the catalog has, to compare. No account needed. Parameters: id (path). - `GET /v1/compare` comparePrintings: Compare two printings word by word (Hebrew-aware). No account needed. Parameters: a (query), b (query). - `POST /v1/units/{id}/translations` addTranslation: Suggest a translation of a unit, as its own text. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/translations/fix` fixTranslation: Suggest a fix to one paragraph of a translation. A token with the write scope, or the site session. Takes a JSON body. - `POST /v1/hanachos/text` addHanachaText: A hanacha's words for a farbrengen or sicha (or a new farbrengen), a paragraph to a segment. A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/texts/{sha256}` getSourceText: A text of a sefer as its source gave it (one chapter or letter, an HTML article). No account needed. Parameters: sha256 (path). - `GET /v1/recordings/{id}/transcript` recordingTranscript: A recording's transcript, with its sync by paragraph and word. No account needed. Parameters: id (path). - `POST /v1/recordings/{id}/transcript/fix` fixTranscript: Fix the words of one paragraph of a transcript (a suggestion). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/recordings/{id}/sync/anchor` anchorSync: The Rebbe is saying this line now: set a paragraph (or word) at atMs, lock it, move what follows. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/recordings/{id}/sync/confirm` confirmSync: The sync is right: mark every paragraph checked. A token with the write scope, or the site session. Parameters: id (path). - `GET /v1/recordings/{id}/hanacha` recordingHanacha: The hanacha synced to this recording, paragraph by paragraph. No account needed. Parameters: id (path). - `GET /v1/files/{sha256}` getFile: A file's size, rights and address, what was made from it, and its page fix. No account needed. Parameters: sha256 (path). - `GET /v1/files/{sha256}/similar` similarFiles: Held files that look like this one (the same scan or recording in other bytes): a machine's guess. No account needed. Parameters: sha256 (path). - `GET /v1/page-fixes/drive/{id}` driveFix: What a PDF on Google Drive needs to read straight, by its Drive id, or the reading copy to open instead. No account needed. Parameters: id (path). - `GET /v1/drive/{id}` driveFile: A Google Drive file the catalog links to (a hanacha's PDF, an Otzros scan), read for the site's reader and player. No account needed. Parameters: id (path). - `GET /objects/{sha256}` getObject: A file's bytes, while its rights let it be served. No account needed. Parameters: sha256 (path). - `GET /manifests/{collection}/{name}` getManifest: A published manifest: reading copies, page fixes (facts about files, open like the catalog). No account needed. Parameters: collection (path), name (path). - `GET /v1/scans/{id}/pages` scanPages: A served scan's page images and thumbnails, and its IIIF manifest. No account needed. Parameters: id (path). - `POST /v1/uploads` upload: Add a file: a recording of a farbrengen; a hanacha's PDF for a farbrengen or sicha; a scan (another scan of a printing, a new printing of a sefer, a teshura); or other material (a new sefer, a letter, a document). A token with the write scope, or the site session. Parameters: what (query), for? (query), eventTitle? (query), eventDate? (query), kind? (query), set? (query), author? (query), genre? (query), unit? (query), rights (query), as? (query), title? (query), publication? (query), publisher? (query), year? (query), printing? (query), families? (query), simcha? (query), date? (query). Takes a JSON body. - `POST /v1/uploads/check` checkUpload: Before an upload: whether we have it (its sha256, a few page hashes) and what it likely is. A token with the write scope, or the site session. Takes a JSON body. - `POST /v1/uploads/propose` proposeUpload: Before adding something new: the machine's guess of what it is and where it belongs, from its name (a date in it, words of a title), and files already held that look like it. A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/files/{sha256}/about` fileAbout: A file's own page: its rights, where it came from, what was made from it, and what uses it. No account needed. Parameters: sha256 (path), limit? (query). - `GET /v1/covers` covers: Sefarim's covers, drawn from their title pages, while their PDFs are served or linked. No account needed. Parameters: ids? (query). - `GET /v1/works/{id}/cover` workCover: A sefer's cover, the page a person chose, and the PDFs (served, or linked) its title page may be chosen from. No account needed. Parameters: id (path). - `POST /v1/entities/{id}/restore` restoreItem: Suggest restoring an earlier version of an item. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `GET /v1/suggestions` listSuggestions: Suggestions, oldest sent first, by status or author; with `state`, the list of conversations, newest first (numbers, reviewers, approvals, the issues each closes) with counts. No account needed. Parameters: status? (query), state? (query), author? (query), reviewer? (query), q? (query), postReview? (query), limit? (query), cursor? (query). Paged (cursor). - `POST /v1/suggestions` createSuggestion: Start a suggestion (a draft): add items to it, then submit it. A token with the write scope, or the site session. Takes a JSON body. - `POST /v1/suggestions/quick` suggestFix: Suggest a fix in one step: a new version of one item, with a few words on why, sent for review. A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/suggestions/{id}` getSuggestion: The review view: each item before and after, clashes with main, and the reviewer's advice (machine-written, `machine: true`). No account needed; signed in, a little more. Parameters: id (path). - `PATCH /v1/suggestions/{id}` editSuggestion: Change the title or description of your suggestion (@mentions and "Fixes #12" are read again). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `GET /v1/suggestions/{id}/conversation` suggestionConversation: A suggestion's timeline (comments, reviews, events), the reviewers asked, and the issues it closes. No account needed; signed in, a little more. Parameters: id (path). - `POST /v1/suggestions/{id}/comments` commentOnSuggestion: Comment on a suggestion, answer a comment, or comment on one field of one item. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/suggestions/{id}/reviews` reviewSuggestion: Review: approve (it goes into the catalog, where you may merge it), request changes (sent back), or comment; with comments on fields. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/suggestions/{id}/review-requests` requestReview: Ask people to review (asking again asks again). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `DELETE /v1/suggestions/{id}/review-requests/{username}` removeReviewRequest: Stop asking someone to review. A token with the write scope, or the site session. Parameters: id (path), username (path). - `PUT /v1/suggestions/{id}/items` putSuggestionItem: Add or change one item in a draft suggestion (data null deletes it). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/suggestions/{id}/submit` submitSuggestion: Send for review (runs the automatic checks). A token with the write scope, or the site session. Parameters: id (path). - `POST /v1/suggestions/{id}/approve` approveSuggestion: Approve and merge (keepers of its sets, stewards). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/suggestions/{id}/send-back` sendBackSuggestion: Send back with a note. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/suggestions/{id}/review-live` reviewLive: Review a live change after it went live: keep it (approve) or undo it (revert). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/suggestions/{id}/withdraw` withdrawSuggestion: Withdraw your suggestion. A token with the write scope, or the site session. Parameters: id (path). - `POST /v1/suggestions/{id}/revert` revertSuggestion: Undo a merged suggestion (a new suggestion that reverses it). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/suggestions/words` suggestWords: A page's words fixed segment by segment: one segment's new words, a segment added after it or taken out, or a page's first words, sent for review. A token with the write scope, or the site session. Takes a JSON body. - `POST /v1/suggestions/contents-map` mapContents: Map pages of a publication to the unit they hold (an existing unit, a new one, or words). A token with the write scope, or the site session. Takes a JSON body. - `POST /v1/comments/{id}/resolve` resolveComment: Resolve (or unresolve) a comment on a suggestion's field. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `GET /v1/tree` catalogTree: The catalog as a tree: the top sets (or one set or sefer), the sets and items under them, and how much each holds. No account needed. Parameters: root? (query), depth? (query), limit? (query). - `POST /v1/organize/preview` previewOrganize: What a plan of moves, renames, orderings, new sets and merges would change, item by item, saved nowhere. A token with the write scope, or the site session. Takes a JSON body. - `POST /v1/organize` organize: Organize the catalog: a plan becomes one suggestion, sent for review (apply: true approves it at once where you may approve it yourself). A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/entities/{id}/talk` itemTalk: An item's talk page: the conversation about it. No account needed. Parameters: id (path). - `POST /v1/entities/{id}/talk` commentOnItem: Comment on an item's talk page. A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/comments/{id}/hide` hideComment: Hide a comment (its author, or a steward). A token with the write scope, or the site session. Parameters: id (path). - `PATCH /v1/comments/{id}` editComment: Change your own comment (on a talk page, a suggestion or an issue). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/reports` report: Report a problem (no account needed: a captcha and an hourly limit instead). No account needed; signed in, a little more. Takes a JSON body. - `GET /v1/reports` listReports: A set's inbox of reports (stewards, and the set's keepers). A token (read scope) or the site session. Parameters: set? (query), status? (query). - `POST /v1/reports/{id}/close` closeReport: Resolve or dismiss a report (keepers). A token with the write scope, or the site session. Parameters: id (path). Takes a JSON body. - `POST /v1/takedowns` requestTakedown: Ask for a file to stop being served (no account needed); stewards answer within two weeks. No account needed. Takes a JSON body. - `POST /v1/teshuros/{id}/family-request` familyRequest: A family's request that a teshura not be shown (no account needed): its scans stop being served at once, and stewards review it. No account needed; signed in, a little more. Parameters: id (path). Takes a JSON body. - `GET /v1/issues` listIssues: Issues, newest first, with open and closed counts; private ones only for those who may read them. No account needed; signed in, a little more. Parameters: state? (query), label? (query), type? (query), set? (query), entity? (query), assignee? (query), author? (query), q? (query), before? (query), limit? (query), cursor? (query). Paged (cursor). - `POST /v1/issues` openIssue: Open an issue (about an item, or the catalog at large). A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/issues/templates` issueTemplates: The kinds of issue and the words each starts with. No account needed. - `GET /v1/issues/{number}` getIssue: An issue, its timeline, what the reader may do, and the suggestions that close it. No account needed; signed in, a little more. Parameters: number (path). - `PATCH /v1/issues/{number}` editIssue: Change its title or words (its author, keepers, stewards). A token with the write scope, or the site session. Parameters: number (path). Takes a JSON body. - `POST /v1/issues/{number}/state` setIssueState: Close as completed or not planned, or reopen. A token with the write scope, or the site session. Parameters: number (path). Takes a JSON body. - `PUT /v1/issues/{number}/labels` setIssueLabels: Set its labels (keepers, stewards, trusted people). A token with the write scope, or the site session. Parameters: number (path). Takes a JSON body. - `PUT /v1/issues/{number}/assignees` setIssueAssignees: Set who it is assigned to (yourself; others when you may triage). A token with the write scope, or the site session. Parameters: number (path). Takes a JSON body. - `POST /v1/issues/{number}/visibility` setIssueVisibility: Make it private or public (stewards and keepers). A token with the write scope, or the site session. Parameters: number (path). Takes a JSON body. - `POST /v1/issues/{number}/comments` commentOnIssue: Comment on an issue, or answer a comment. A token with the write scope, or the site session. Parameters: number (path). Takes a JSON body. - `GET /v1/labels` listLabels: Every label and how many open issues carry it. No account needed. - `POST /v1/labels` createLabel: Make a label (stewards). A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/people` searchPeople: People to @mention: handles that start with, or names that contain, what is typed; those in the conversation first. No account needed. Parameters: q? (query), ids? (query), thread? (query), limit? (query). - `GET /v1/people/{username}` getProfile: A person's public page: who they are, their counts and recent activity (an old handle finds them too, with `movedFrom`). No account needed. Parameters: username (path), limit? (query). - `GET /v1/threads` searchThreads: Suggestions and issues to #mention, by number or words. No account needed; signed in, a little more. Parameters: q? (query), limit? (query). - `GET /v1/threads/{number}` threadByNumber: Which of the two #12 is: a suggestion or an issue, and its id. No account needed; signed in, a little more. Parameters: number (path). - `GET /v1/inbox` inbox: Your inbox, newest first: mentions, review requests, assignments and what you follow. A token (read scope) or the site session. Parameters: filter? (query), before? (query), limit? (query), cursor? (query). Paged (cursor). - `GET /v1/inbox/count` inboxCount: How many inbox lines are unread. A token (read scope) or the site session. - `POST /v1/inbox/read` markInboxRead: Mark inbox lines read (or unread): by id, by conversation, or all. A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/missing` missing: The Missing board: farbrengens without a recording or a text, sefarim without a scan, files lost upstream. No account needed. Parameters: kind (query), within? (query), limit? (query). - `GET /v1/projects` listProjects: Projects working through a gap, with their progress. No account needed. Parameters: status? (query). - `POST /v1/projects` createProject: Open a project on a gap (farbrengens without recordings or texts, recordings to sync, pages to proofread). A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/projects/{slug}` getProject: A project, its progress and what is left to do. No account needed. Parameters: slug (path). - `POST /v1/projects/{slug}/next` claimNext: Hand me the next item nobody holds (held for you for a few hours). A token with the write scope, or the site session. Parameters: slug (path). - `POST /v1/projects/{slug}/release` releaseClaim: Let go of an item you held. A token with the write scope, or the site session. Parameters: slug (path). Takes a JSON body. - `POST /v1/projects/{slug}/close` closeProject: Close a project (its keepers, stewards). A token with the write scope, or the site session. Parameters: slug (path). - `GET /v1/follows` listFollows: What you follow, the items themselves, and what changed in them lately. A token (read scope) or the site session. Parameters: limit? (query). - `POST /v1/follows` follow: Follow or unfollow an item, set, project, suggestion or issue. A token with the write scope, or the site session. Takes a JSON body. - `GET /v1/places` listPlaces: Where you stopped reading and listening lately (never cached). A token (read scope) or the site session. Parameters: kind? (query), key? (query), limit? (query). - `PUT /v1/places` savePlace: Keep where you stopped in one thing. A token with the write scope, or the site session. Takes a JSON body. - `DELETE /v1/places` forgetPlace: Forget one place. A token with the write scope, or the site session. Parameters: kind (query), key (query). - `GET /v1/webhooks` listWebhooks: Your webhooks: addresses every merge is posted to. A token (read scope) or the site session. - `POST /v1/webhooks` createWebhook: Add a webhook (up to five); its signing secret is shown this once. A token with the write scope, or the site session. Takes a JSON body. - `DELETE /v1/webhooks/{id}` deleteWebhook: Remove a webhook. A token with the write scope, or the site session. Parameters: id (path). - `GET /.well-known/oauth-protected-resource` protectedResource: The API's Protected Resource Metadata (RFC 9728): which authorization server gives its tokens. No account needed. - `GET /.well-known/oauth-protected-resource/mcp` mcpProtectedResource: The MCP server's Protected Resource Metadata (RFC 9728), named in its 401's WWW-Authenticate. No account needed. - `GET /.well-known/oauth-authorization-server` authorizationServer: Authorization Server Metadata (RFC 8414): the endpoints, scopes read and write, PKCE S256, registration and Client ID Metadata Documents. No account needed. - `POST /oauth/register` oauthRegister: Register an app (RFC 7591): its name and redirect addresses; a secret only if it asks for one. No account needed. Takes a JSON body. - `GET /oauth/authorize` oauthAuthorize: Start connecting (authorization code with PKCE): the person is sent to the site's consent page, then back to the app. No account needed. Parameters: response_type (query), client_id (query), redirect_uri? (query), scope? (query), state? (query), code_challenge (query), code_challenge_method (query), resource? (query), ui_locales? (query). - `POST /oauth/token` oauthToken: Trade a code (with its PKCE verifier) or a refresh token for an access token (an hour) and a new refresh token. No account needed. Takes a JSON body. - `POST /oauth/revoke` oauthRevoke: Revoke an access or refresh token (RFC 7009): the whole connection ends. No account needed. Takes a JSON body. - `GET /v1/mirrors` mirrors: Everything a mirror needs: the git mirror, the release keys, every edition and its dumps. No account needed. - `GET /v1/editions` editions: Catalog editions (dated snapshots) and their dumps, each with its size, sha256 and address. No account needed. - `GET /v1/editions/{tag}/manifest.json` editionManifest: An edition's signed manifest (Ed25519), exactly as signed. No account needed. Parameters: tag (path). - `GET /v1/editions/{tag}/SHA256SUMS` editionChecksums: An edition's checksums, for sha256sum -c. No account needed. Parameters: tag (path). - `GET /dumps/{tag}/{name}` getDump: One of an edition's dumps (SQLite, JSON Lines, Parquet). No account needed. Parameters: tag (path), name (path). - `GET /manifests/iiif/{file}` iiifManifest: A served scan as a IIIF Presentation 3 manifest, for any IIIF viewer. No account needed. Parameters: file (path). - `GET /oai` oai: OAI-PMH 2.0 for libraries (oai_dc records), when switched on. No account needed. Parameters: verb (query), metadataPrefix? (query), identifier? (query), from? (query), until? (query), set? (query), resumptionToken? (query). - `POST /oai` oaiPost: OAI-PMH, the same arguments sent as a form. No account needed. Takes a JSON body. - `GET /llms.txt` llmsTxt: A short guide for AI agents (llms.txt). No account needed. - `POST /mcp` mcp: The Model Context Protocol server (Streamable HTTP, JSON answers, no sessions). No account needed; signed in, a little more. Takes a JSON body. - `GET /mcp` mcpStream: Not offered: this server opens no event stream. No account needed. - `DELETE /mcp` mcpEnd: Not offered: there are no sessions to end. No account needed. - `POST /v1/auth/email/unsubscribe` unsubscribe: Stop email updates, from the link in any of them (no sign-in). No account needed. Parameters: token? (query). Takes a JSON body.