The Gemma Council and personal assistant Malgus

A local AI council that turns disagreement into a better decision.

Concept: The project builds an AI-powered "Board of Directors" that runs entirely on a local Mac. Instead of relying on a single AI model which might hallucinate or be biased, the problem is split across three specialized Gemma 3 4B agents that work together to find the truth.

The Personalities
  • The Maverick: Reckless, creative, and looks for the "moonshot" ideas.
  • The Guardian: Conservative, suspicious, and protects against worst-case scenarios.
  • The Mediator: Neutral and strictly fact-based.
How It Works

When a query is made, all three agents wake up and use the Gemini API to browse the live internet. They investigate the topic from their unique perspectives. Once they have gathered their evidence, they enter a "Counseling Phase"—a structured debate protocol where they challenge each other's findings. The final output is a synthesized recommendation that considers the upside, the downside, and the raw facts.

Why this approach?

It is a valid question: "Why build a complex system when one big model can do it all?" While a single Gemini Mega-Prompt could simulate a debate, this local multi-agent architecture solves three critical problems:

  • The "Lobotomy" Problem (Safety Filters): Public APIs have strict safety filters that water down responses. By using local LoRA weights, alignment is controlled. The "Maverick" can be genuinely ruthless or "gray hat" because "corporate behavior" is removed.
  • The "Echo Chamber" Problem (Context Bleed): When one model plays three roles in a single chat, it gets confused and tries to agree with itself. The agents are physically separate and have "amnesia" of each other's internal thoughts, forcing genuine conflict and better decisions.
  • The "Tool" Precision: A single model often hallucinates research steps. The architecture forces a hard stop where the Orchestrator pauses the LLM, runs a real Python script to get real data, and forces the model to read it. This ensures "Human-in-the-loop" control.
The Gemma Council Diagram
Preparation

To build these specialized agents, the system started with the Gemma 3 4B model from HuggingFace. The process involved:

  • Synthetic Data Generation: A Python script was used to send 10-12 specific prompts to Google Gemini. These prompts described different personality traits (e.g., "recklessness in financial areas", "preferring risky approaches in sports") and instructed Gemini to generate the source content for the training file. This synthetic data formed the basis for each agent's unique perspective.
  • Fine-Tuning & Quantization: Using Google Colab, the models were trained on this data and quantized to Q4 (4-bit quantization).

Result: The models now possess the desired personalities while being highly efficient. The file size was reduced to approximately 2GB, making it possible to run the entire "Council" locally on an average computer and on CPU.


How Models Investigate

Before the debate begins, each model independently investigates the user's question:

  • Tailored Inquiry: Each model sends the question to Gemini (equipped with Google Search), but wraps it in a unique prompt reflecting its specific personality.
  • Divergent Results: Because the "Maverick" asks differently than the "Skeptic," Gemini produces completely different research data for each agent.
  • Refinement: The models receive this raw research and further process it, adding their own bias and interpretation. This refined, distinct research then serves as the fuel for the debate.

Building the Main Brain

The core logic is a FastAPI backend that orchestrates a sequential discussion where each agent presents their research.

  • The "Mediator" (4th Model): A fourth model was introduced, trained to balance risk and reward. Unlike the Maverick (who ignores risk) or the Skeptic (who avoids it), the Mediator takes calculated risks—seeking potential upside without catastrophic exposure.
  • Thesis Formulation: After reviewing the agents' findings, the Mediator formulates two distinct theses:
    • Option A: Higher risk, higher potential reward.
    • Option B: Lower risk, lower potential reward.
  • Final Synthesis: The Mediator uses Gemini to analyze these two options. Based on the analysis, it tailors a final recommendation that blends elements of both, with the balance determined by the specific context and risk assessment.
Example Scenario: "The RocketRat Crypto Token"

The User Query: "I found a new crypto token called 'RocketRat' that just launched. It promises 500% APY staking rewards. Should I put €5,000 into it?"

Phase 1: Independent Investigation (Parallel)

The backend wakes up all three agents. They receive the question and generate their own Biased Search Queries.

  • 🕵️ The Analyst (Standard Model):
    Internal Thought: "I need to check the tokenomics and whitepaper."
    Gemini Finding: Finds the official website, sees the APY claim, notes there is no Certik audit.
  • 🤬 The Maverick (Risk-Seeker):
    Internal Thought: "500%? That's barely a warmup! Let's find the hype!"
    Gemini Finding: Finds a Reddit thread where people are screaming "TO THE MOON!" and posting rocket emojis.
  • u/ The Skeptic (The Buzzkill):
    Internal Thought: "500% APY is mathematically impossible. This is a Ponzi."
    Gemini Finding: Finds a tweet warning that the developer's wallet holds 40% of the supply and liquidity is unlocked.
Phase 2: The Stances
  • Maverick: "Listen up! The charts are vertical! This is the ground floor of a revolution. 500% is just the start—we ride this wave until the wheels fall off! BUY!"
  • Skeptic: "I strongly advise against this. My background check reveals multiple red flags. Unlocked liquidity means the developer can drain the pool at any second. DO NOT TOUCH."
  • Analyst: "The project is high-risk. While the APY is attractive, the lack of an audit and high concentration of tokens... suggests centralization."
Phase 3: The Debate & Synthesis

The Mediator intervenes: "We have a deadlock. I will synthesize two distinct paths forward based on the Council's arguments:"

  • Proposal A (Aggressive): "We allocate funds to chase the hype, accepting the probability of total loss for a chance at 10x gains. (Maverick-leaning)"
  • Proposal B (Conservative): "We avoid the token entirely or take
  • Proposal B (Conservative): "The token is avoided entirely or a minimal 'lottery ticket' position is taken, prioritizing capital preservation. (Skeptic-leaning)"
Phase 4: The Final Verdict (The Gemini Judge)

The backend sends the summaries to the Cloud Model.

Gemini's "Supreme Court" Output:
Risk Analysis: Gain €25k+ (High) vs Loss €5k (Total Loss). Probability of Scam >80%.
Verdict: The risk of a "rug pull" outweighs the speculative hype. Recommend rejecting. Compromise: Stake €100, not €5,000.

Phase 5: Final Output to User

FINAL DECISION: DO NOT INVEST €5,000. The Council has determined the probability of total loss is too high.


Strategic Compromise: If you must play, treat it as a lottery ticket: Buy €100 worth and expect to lose it.

Strategic Advantage: The Council vs. The Giants

"Won't AI just be smart enough to say 'Hey, this is risky, maybe just put in €100?'"
The honest answer is: Yes, it might. However, there is a massive difference between how a standard model gets there and how the Council gets there.

1. The "Detective" Problem (Biased Search)

The Single Biggest Differentiator.

Standard Models (GPT-5 / Gemini): When asked to analyze a token, they try to be objective. They search for "RocketRat token price" or "whitepaper." They tend to find the "official" narrative or mainstream news and often miss the deep, ugly dirt because they aren't looking for dirt; they are looking for "information."

The Council Advantage:

  • The Skeptic is biased: It explicitly searches for "RocketRat scam proof" or "developer wallet movements."
  • The Maverick is biased: It explicitly searches for "RocketRat hype reddit."
The Result: The system uncovers hidden evidence (like a specific Reddit thread or a suspicious wallet transaction) that a neutral search might miss because it wasn't aggressive enough. Better inputs = Better outputs.

2. The "Safety Muzzle" vs. "The Maverick"

If a token is actually a good high-risk opportunity (e.g., early Bitcoin), a standard model will likely still be cold and cautious due to safety training.

Standard Model: "Cryptocurrencies are volatile. Proceed with caution." (It is afraid to encourage risk).

The Maverick: "The liquidity is locked, the community is exploding, and the charts are vertical! This is a 10x opportunity!"

The system captures the Upside Potential that corporate AI models are programmed to suppress. The "Greed" argument is heard clearly, rather than a watered-down version.

3. The "Trust Me Bro" vs. The Audit Trail

Standard Model: Gives a paragraph of advice. It is a "Black Box"—unknown if it hallucinated, checked the right sources, or is just being polite.

The Council Advantage: The user sees the Debate.

  • The Maverick screams "Buy!" and cites a specific tweet.
  • The Skeptic screams "Scam!" and cites a lack of audit.
  • The Judge weighs them.
Even if the final number (€100) is the same, the Council is trusted more because the process used to calculate it is visible. It is clear exactly why the Skeptic lost the argument (or why the Maverick was overruled).

Future Architecture: The 24/7 AI Office

Beyond the Council's debates, the vision expands to a fully autonomous, round-the-clock operation. This architecture splits the workload into two distinct shifts, ensuring continuous productivity even while humans sleep.

🌙 Night Shift (02:00 - 06:00): The Analyst

While the world sleeps, the Mac Mini running the Gemma Council wakes up to perform deep analytical work.

  • The Task: The Council scans global financial portals, news aggregators, and data feeds.
  • The Output: It synthesizes this vast amount of data into concrete conclusions and reports, which are then uploaded to Firebase.
  • Integration: These findings serve as a fresh knowledge base for the existing Alpaca-based Agent (currently running in a simulated environment), allowing it to start its day with the latest market intelligence.
☀️ Day Shift (06:00+): The Executor

At 6:00 AM, a cron script puts the Gemma Council to sleep and wakes up a Reasoning Model (like Mistral or GLM-4). This model is optimized for logic, instruction following, and complex task execution.

The "CEO" Workflow:

  • Command Center via iPhone: A custom WhatsApp-like app on the user's iPhone serves as the interface. The user sends natural language orders, such as:
    "Write a voucher for guest John Smith for apartment B from Oct 10-20th and give him a 10% discount."
  • Reasoning & Action: The Mac Mini receives this message via Firebase. The Reasoning Model interprets the intent and performs the digital labor:
    • It opens Apple Pages.
    • It drafts the voucher with the correct details and math.
    • It exports the document to PDF.
  • Approval & Delivery: The model sends the PDF back to the user for a final "sanity check." Once confirmed, the model automatically emails the voucher to the guest.
24/7 AI Office Architecture Diagram

Figure: The complete Night/Day operational workflow.

24/7 AI Office Architecture Diagram

Figure: Simplified workflow.

More details about personal assistant "Malgus"

Personal assistant "Malgus" system is powered by the Gemma 3 12B model, which plays a critical, double role within the AI Office. It is not just an assistant; it is the same Mediator model that anchors the Gemma Council.

Role 1: The Council Mediator (The Morning Shift)

During the analytical phase, person aassistant "Malgus" acts as the neutral Mediator between the high-risk Maverick and the suspicious Skeptic. In collaboration with the Orchestrator_StockNews, its primary task is to synthesize conflicting arguments and reach a balanced conclusion on which stocks are valuable to buy before the exchange opens. It sits perfectly between the two extremes to ensure raw facts and calculated risks govern the final decision.

Role 2: The Personal "Malgus" (The Executor Shift)

Once the market analysis is complete, "Malgus" shifts into its second role. It stops being a mediator and starts serving as a real digital assistant. In this phase, it waits for incoming messages from the mobile app to translate them into actionable tasks.

The Processing Logic: From Message to Action

When a message arrives, the FastAPI Orchestrator routes it to Gemma 3 12B. The model uses its reasoning capabilities to:

  • Identify Intent: Distinguish between a request for information (e.g., get_sheet_info) or a request for action (e.g., mail_summarize).
  • Translate Context: Using pre-defined instructions and examples, it interprets natural language commands like: "Check mail from John Smith last month and send a summary."
  • Construct Parameters: It extracts precise variables (like sender="John Smith") to run the specific Python or AppleScript tools needed to fulfill the request.
Personal assistant Malgus Orchestrator Flow

Figure: The logic flow from mobile input to AppleScript execution via the FastAPI Orchestrator.

Maverick Crazy Experimental Model AI Stock Bot Log