<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>saidrassai</title><link>https://saidrassai.github.io/</link><description>Recent content on saidrassai</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 15 Jun 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://saidrassai.github.io/index.xml" rel="self" type="application/rss+xml"/><item><title>Fine-tuning LFM2.5-1.2B-Instruct with GRPO</title><link>https://saidrassai.github.io/blog/finetune-lmf2-5/</link><pubDate>Mon, 15 Jun 2026 00:00:00 +0000</pubDate><guid>https://saidrassai.github.io/blog/finetune-lmf2-5/</guid><description>In this notebook, we will explore the core concepts of GRPO (Group Relative Policy Optimization) by fine-tuning LFM2.5-1.2B-Instruct using Unsloth .
GRPO is a reinforcement learning algorithm designed for training language models with reward signals instead of labeled examples. In contrast to supervised fine-tuning (SFT), where you tell the model the exact right answer, GRPO lets the model explore different outputs and reinforces the ones that score higher on the reward functions.</description></item><item><title>Building an ENTERPRISE CPU RAG RESEARCH 2025–2026</title><link>https://saidrassai.github.io/blog/entreprise-cpu-rag/</link><pubDate>Fri, 05 Jun 2026 00:00:00 +0000</pubDate><guid>https://saidrassai.github.io/blog/entreprise-cpu-rag/</guid><description>This post documents the end-to-end thinking behind reproducing a Fin-R1-style financial reasoning dataset, and what that implies for building a finance agent that can answer questions with genuine chain-of-thought reasoning instead of backward-engineered justifications.
1. HYBRID RETRIEVAL: ARCHITECTURE, FUSION &amp;amp; OPTIMIZATION 1.1 Key Papers &amp;amp; Industry Systems (2025–2026) Paper / System Venue / Org Key Finding for Production RAG Hybrid Search (BM25 + Dense) Industry Standard (Elastic, Vespa, Weaviate 2025) RRF (Reciprocal Rank Fusion) k=60 remains SOTA for zero-shot.</description></item><item><title>Building a Finance Agent and Dataset: From Research Note to Replication</title><link>https://saidrassai.github.io/blog/finance-agent-journey/</link><pubDate>Tue, 02 Jun 2026 00:00:00 +0000</pubDate><guid>https://saidrassai.github.io/blog/finance-agent-journey/</guid><description>This post documents the end-to-end thinking behind reproducing a Fin-R1-style financial reasoning dataset, and what that implies for building a finance agent that can answer questions with genuine chain-of-thought reasoning instead of backward-engineered justifications.
Why replicate Fin-R1-Data? The official Fin-R1-Data dataset is not publicly available. The authors at SUFE-AIFLM-Lab promised a March 2025 release but never followed through. That leaves practitioners with three options:
Reconstruct it from the raw open sources using the described two-stage pipeline.</description></item></channel></rss>