← Registry

AI & Machine Learning

bakedin.co

Processes and analyzes academic papers to extract structured claims, compute evidence grades, and build training corpora for language models.

1 endpoint35 known toolsCached registry data

ENDPOINT 1

https://pantry.bakedin.co/corpus/mcp

No auth detected

Known tools 35

add_to_corpus

Mark license-verified papers as included in the training corpus.

analyze_papers

Extract PICO-structured claims from papers using Bedrock Haiku, then compute GRADE-lite evidence grades deterministically.

assemble_corpus

Build a CPT training corpus JSONL from extracted papers.

classify_figures

Run Bedrock Haiku classification on unclassified figures.

corpus_candidates

Papers that are license-verified and ready for corpus inclusion.

corpus_chunk_search

CHUNK-level semantic search — returns the actual paragraphs of body text most relevant to your query, with their paper context.

corpus_gap_report

Gap analysis and finish-line cost projection.

corpus_hybrid_search

BEST general-purpose search: fuses corpus_keyword_search + corpus_semantic_search via Reciprocal Rank Fusion (RRF).

corpus_ingest_graph_edges_jsonl

Ingest course_content_graph edges from S3-hosted JSONL.

corpus_ingest_jsonl

Ingest staged papers + chunks from S3-hosted JSONL into corpus_papers + corpus_section_chunks.

corpus_keyword_search

Full-text keyword search over the corpus (172K papers).

corpus_semantic_search

Meaning-based search over the corpus via pgvector cosine similarity.

corpus_state

Current corpus state across all aspects: paper intake, pipeline stages, derived artifacts (claims/SFT/figures/safety/cards), full-text sections, books, regulatory/tribal/underwriting content, quality and relevance distributions, and empirical cost history.

corpus_state_trend

Compares the latest corpus snapshot to one from N days ago, showing deltas on headline metrics (papers, abstracts, license-verified, CPT tokens, claims, SFT pairs + source papers, paper sections, books).

datacenter_anatomy

Assemble the cross-axis 'Anatomy of a Datacenter': a facility's scale + siting (infra_projects), the energy build-out in its state, the subsidies/cost-shift in its jurisdiction (infra_subsidies), and the corpus papers explaining the underlying constraints.

deduplicate

Find and report duplicate papers across different sources (matching DOI or ArXiv ID).

discover_domain

Fast metadata-only indexing.

domain_coverage

Stats on paper coverage per domain and overall progress toward the 300+ paper corpus target.

extract_figures

Download PDFs and extract figures for papers.

extract_full_text

Download PDF, extract full text with section parsing (IMRaD), chunk for training.

extract_text

Extract text from papers that have abstracts or PDF URLs.

figure_catalog

Search and browse extracted figures with filters.

figure_detail

Get full metadata for a specific figure by figure_id.

generate_sft_from_claims

Generate SFT training Q&A pairs from extracted claims using Bedrock Haiku.

harvest_all

Run full harvest across all domains and all sources.

harvest_domain

Harvest papers for a specific domain from a source API.

ingest_queue_health

Queue-state invariants for the ingest pipeline.

paper_detail

Full details for a single paper by paper_id (e.g.

paper_registry

Query the paper registry with optional filters.

processing_status

Pipeline processing progress across all domains and stages.

promote_papers

Verify licenses and promote high-quality discovered papers.

reject_paper

Mark a paper as rejected with a reason.

rejection_log

View the audit log of papers rejected during harvesting.

run_canary

Run canary test: preflight checks, process 5 papers through full pipeline, run all validators.

verify_licenses

Batch-verify license status for papers in 'discovered' status.