د درسهای عمیقتر فقط در نسخههای ابری
نسخههای ابری دوره چهار درس اضافی دارند که در نسخه پایه حذف شدهاند. علت: مشتریان ابری اغلب RAG-heavy هستند (Bedrock KB، Vertex Search) و agent-heavy (Bedrock Agents، Vertex Reasoning Engine)، پس برنامه آموزشی نیاز به تقویت روی کیفیت retrieval و orchestration ابزار دارد.
۶.د.۱ Reranking
هدف یادگیری: درک اینکه چرا vector search تنها کافی نیست و reranking کجا وارد pipeline میشود.
چرا وجود دارد؟ vector search تنها به سقف میخورد — embeddingها مطابقتهای دقیق lexical را میاندازند بیرون (یک شماره مدل، یک کد محصول، نام خانوادگی یک شخص). hybrid search (vector + BM25) بخشی از این زمین را برمیگرداند، اما همچنان مجبورید over-retrieve کنید تا recall را حفظ کنید. یک reranker، یک cross-encoder است که جفت (query, candidate) را بهصورت مشترک امتیازدهی میکند و این اشتباه over-retrieval را اصلاح میکند.
Pipeline:
- top-N (معمولاً ۱۰۰ تا ۱۵۰) candidate را با هر روش recall-oriented که دارید (vector، BM25، یا hybrid) بازیابی کنید.
- جفتهای
(query, candidate_text)را به یک reranker بفرستید (Coherererank-english-v3.0، Voyagererank-2، BGE reranker، و غیره). - top-K (۵ تا ۲۰) candidate رتبهبندیشده را بهعنوان context نهایی RAG برگردانید.
Trade-offs:
- ۱۰۰ تا ۲۰۰ میلیثانیه latency اضافه میکند (یک HTTPS call به reranker، روی batchها قابل موازیسازی).
- ~$0.002 بهازای هر query با قیمت Cohere اضافه میکند.
- در benchmark contextual-retrieval Anthropic، reranking مقدار Pass@10 را از ۸۷.۱۵٪ (پایه) به ۹۵.۲۶٪ رساند، و وقتی روی contextual embeddings + BM25 stack شد، نرخ شکست retrieval در top-20 را ۶۷٪ کاهش داد.
Python مرجع (از cookbook Anthropic):
import cohere
co = cohere.Client(os.getenv("COHERE_API_KEY"))
# Step 1: over-retrieve from your vector DB
semantic_results = db.search(query, k=k * 10)
documents = [
f"{r['metadata']['original_content']}\n\nContext: {r['metadata']['contextualized_content']}"
for r in semantic_results
]
# Step 2: rerank
rerank_response = co.rerank(
model="rerank-english-v3.0",
query=query,
documents=documents,
top_n=k,
)
# Step 3: keep only the top-k reranked docs
final_docs = [
{"chunk": semantic_results[r.index]["metadata"], "score": r.relevance_score}
for r in rerank_response.results
]
در نسخه AWS، این درس به Bedrock Knowledge Bases متصل میشود (که Cohere Rerank را بهصورت native بهعنوان یک rerankingConfiguration روی RetrieveAndGenerate ادغام میکند). در نسخه GCP، به Vertex AI Ranking API یا semantic ranker داخلی Vertex AI Search متصل میشود.
۶.د.۲ Contextual Retrieval — روش کامل Anthropic
هدف یادگیری: یاد گرفتن روش Anthropic که نرخ شکست retrieval را تا ۶۷٪ کاهش میدهد.
مساله. یک chunk RAG ساده، context سندی خود را گم میکند. این chunk «درآمد شرکت ۳٪ نسبت به سهماهه قبل رشد کرد» به شکل قابلاستفادهای embed نمیشود — هیچ سرنخی نیست که کدام شرکت یا کدام سهماهه.
راهحل. قبل از embedding، یک جمله context تولید-شده توسط مدل (۵۰ تا ۱۰۰ token) به ابتدای chunk اضافه کنید که آن را داخل سند والد جای دهد. از prompt caching استفاده کنید تا هزینه per-chunk تقریباً صفر بماند (سند، prefix کششده است).
Prompt مرجع تولید context:
<document>
{{WHOLE_DOCUMENT}}
</document>
Here is the chunk we want to situate within the whole document
<chunk>
{{CHUNK_CONTENT}}
</chunk>
Please give a short succinct context to situate this chunk within the
overall document for the purposes of improving search retrieval of the
chunk. Answer only with the succinct context and nothing else.
پیادهسازی Cookbook:
DOCUMENT_CONTEXT_PROMPT = """
<document>
{doc_content}
</document>
"""
CHUNK_CONTEXT_PROMPT = """
Here is the chunk we want to situate within the whole document
<chunk>
{chunk_content}
</chunk>
Please give a short succinct context to situate this chunk within the
overall document for the purposes of improving search retrieval of the
chunk. Answer only with the succinct context and nothing else.
"""
def situate_context(doc: str, chunk: str) -> str:
response = client.messages.create(
model="claude-haiku-4-5", # cheap + fast for this task
max_tokens=1024,
temperature=0.0,
messages=[{
"role": "user",
"content": [
{"type": "text",
"text": DOCUMENT_CONTEXT_PROMPT.format(doc_content=doc),
"cache_control": {"type": "ephemeral"}},
{"type": "text",
"text": CHUNK_CONTEXT_PROMPT.format(chunk_content=chunk)},
],
}],
)
return response.content[0].text
# At index time:
text_to_embed = f"{situate_context(doc, chunk)}\n\n{chunk}"
embedding = voyage_client.embed([text_to_embed], model="voyage-2").embeddings[0]
دستور کامل Anthropic = Contextual Embeddings + Contextual BM25 + Reranking:
- Contextual Embeddings. هر chunk همراه با context پیوستشدهاش embed میشود.
- Contextual BM25. همان chunk contextualized در Elasticsearch (analyser انگلیسی، similarity BM25) برای recall lexical ایندکس میشود.
- Hybrid retrieval. هر دو را اجرا کنید، از هرکدام top 150 را بگیرید، با reciprocal-rank وزندار ادغام کنید (پیشفرض ۰.۸ semantic، ۰.۲ BM25)، dedupe کنید.
- Rerank. union را به Cohere
rerank-english-v3.0بفرستید،top_n=20. - Generate. top-20 را به Claude بهعنوان context RAG بدهید.
کد score-fusion (از cookbook):
for chunk_id in chunk_ids:
score = 0
if chunk_id in ranked_chunk_ids:
i = ranked_chunk_ids.index(chunk_id)
score += semantic_weight * (1 / (i + 1))
if chunk_id in ranked_bm25_chunk_ids:
i = ranked_bm25_chunk_ids.index(chunk_id)
score += bm25_weight * (1 / (i + 1))
chunk_id_to_score[chunk_id] = score
نتایج ablation (benchmark منتشرشده Anthropic):
| روش | Pass@5 | Pass@10 | Pass@20 | کاهش نرخ شکست retrieval |
|---|---|---|---|---|
| RAG پایه (فقط vector) | 80.92 % | 87.15 % | 90.06 % | – |
| + Contextual Embeddings | 88.12 % | 92.34 % | 94.29 % | -35 % |
| + Hybrid Search (BM25) | 86.43 % | 93.21 % | 94.99 % | -49 % |
| + Reranking | 92.15 % | 95.26 % | 97.45 % | -67 % |
هزینه. با prompt caching، تولید chunkهای contextualized حدوداً $1.02 بهازای ۱M token سند هزینه دارد (سند، prefix کششده است؛ فقط chunk نرخ کامل per-call را میپردازد). Anthropic ۶۱.۸۳٪ cache-hit روی input tokens اندازه گرفت.
قاببندی ابری. روی Bedrock، این درس مقابل گزینه «advanced parsing» Bedrock Knowledge Bases + toggle Cohere Rerank آموزش داده میشود، که این pipeline را با یک UX مدیریتشده تقریب میزند. روی Vertex، مقابل document-context augmentation داخلی Vertex AI Search + Ranking API.
۶.د.۳ Batch tool use — فراخوانی موازی toolها در یک turn
هدف یادگیری: ساخت یک meta-tool که چندین tool دیگر را بهصورت همروند اجرا میکند.
چرا وجود دارد؟ Claude میتواند چندین بلوک tool_use در یک پاسخ منتشر کند، اما مشتریان معمولاً به دو حالت شکست میخورند: (الف) نسلهای قدیمی Sonnet در برنامهریزی فراخوانیهای موازی محتاطاند، و (ب) round-trip سریالی toolها از طریق SDK کند است. راهحل، یک meta-tool است — مثلاً به اسم batch_tool — که تنها وظیفهاش بستهبندی N invocation از سایر toolهاست تا توسط harness بهصورت همزمان اجرا شوند.
Schema ابزار:
batch_tool_schema = {
"name": "batch_tool",
"description": (
"Invoke multiple other tools in parallel. Use whenever two or more "
"independent tool calls are needed; they will all run concurrently."
),
"input_schema": {
"type": "object",
"properties": {
"invocations": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"arguments": {"type": "string",
"description": "JSON-encoded arguments"},
},
"required": ["name", "arguments"],
},
}
},
"required": ["invocations"],
},
}
حلقه harness (گزیده):
import asyncio, json
async def execute_batch(tool_use, registered_tools):
invocations = tool_use["input"]["invocations"]
coros = [
registered_tools[inv["name"]](**json.loads(inv["arguments"]))
for inv in invocations
]
results = await asyncio.gather(*coros, return_exceptions=True)
return [
{"tool_name": inv["name"],
"result": (str(r) if not isinstance(r, Exception) else f"ERROR: {r}")}
for inv, r in zip(invocations, results)
]
# Inside the agent loop:
if tool_use["name"] == "batch_tool":
batched_results = await execute_batch(tool_use, registered_tools)
follow_up = client.messages.create(
model=MODEL,
max_tokens=2048,
tools=all_tools + [batch_tool_schema],
messages=history + [
{"role": "assistant", "content": prev_response.content},
{"role": "user", "content": [{
"type": "tool_result",
"tool_use_id": tool_use["id"],
"content": json.dumps(batched_results),
}]},
],
)
قاببندی ابری. روی Bedrock Agents، همین ایده بهصورت بومی به شکل invocation موازی action-group پیاده میشود. روی Vertex Reasoning Engine، فراخوانیهای موازی tool توسط executor engine زمانبندی میشوند. این درس الگوی خام (بالا) را آموزش میدهد تا دانشجو بفهمد runtimeهای مدیریتشده زیر کاپوت چه میکنند.
۶.د.۴ Automated debugging — agentهای خودترمیم
هدف یادگیری: ساخت یک حلقه exec → observe → diff → patch با Claude بهعنوان planner و critic.
الگو. کوچکترین حلقه ممکن «اجرا کن → مشاهده کن → با انتظار مقایسه کن → patch بزن» را بسازید، با Claude در دو نقش planner و critic. این درس در سه لایه ارائه میشود:
- Run-and-observe. یک tool به اسم
run_tests()یاexec_python()که stdout/stderr/exit-code را برمیگرداند. به Claude گفته میشود آن را صدا بزند و خلاصه شکست را بدهد. - Hypothesise. Claude یک کوچکترین edit پیشنهاد میکند (با یک tool به سبک
str_replace_editor)، سپس دوباره اجرا میکند. - Self-critique. یک prompt دوم از Claude میخواهد تلاش کند fix خودش را رد کند — یعنی برای حالتهایی استدلال کند که patch هنوز شکست میخورد. اگر نتوانست رد کند، fix را قبول کنید؛ وگرنه iterate.
اسکلت حلقه مرجع:
def auto_debug(failing_test, max_iters=5):
history = [
{"role": "user", "content":
f"You are a debugging agent. The test below is failing.\n"
f"Loop: (a) call run_tests, (b) read the failure, (c) propose ONE "
f"minimal patch via str_replace_editor, (d) call run_tests again. "
f"After each successful fix, call critique_self to argue why your "
f"fix could still be wrong; iterate until critique_self returns "
f"'no remaining objections'. Failing test: {failing_test}"},
]
for _ in range(max_iters):
resp = client.messages.create(model=MODEL, max_tokens=4096,
tools=DEBUG_TOOLS, messages=history)
history.append({"role": "assistant", "content": resp.content})
if resp.stop_reason == "end_turn": return history
for block in resp.content:
if block.type == "tool_use":
tool_result = TOOLS[block.name](**block.input)
history.append({"role": "user", "content": [{
"type": "tool_result",
"tool_use_id": block.id,
"content": tool_result,
}]})
return history
گام self-critique («تلاش کن fix خودت را رد کنی») همان ترفندی است که Anthropic داخل قابلیت security-review در Claude Code استفاده میکند — به مدل گفته میشود به گزارش خودش حمله کند تا false positiveها فیلتر شوند.
قاببندی ابری. Bedrock Agents یک trace + retry hook داخلی برای این کار افشا میکند؛ Vertex Reasoning Engine یک تنظیم retry-with-reflection مشابه دارد. درس، الگوی زیربنایی را میآموزد تا بین runtimeها قابل پورت باشد.