Recently, I have been investing quite some time building some features for my personal project RecAtlas, and I had one challenge in my hands regarding the standardization of the categories of the activities. At first, I was thinking about implementing some FuzzySearch with PostgreSQL and create some pre-defined categories, and then, a data migration, and a categorization class for the new activities fetched from the API, but then, I realized: "Why not use one of the things that AI does best, which is categorization of information?".

When people think about integrating LLMs into their Rails apps, they usually jump straight to building chat interfaces. But here's the thing - some of the most powerful applications of LLMs are way more subtle. Let's talk about using RubyLLM to solve everyday problems like categorization, intelligent search, and data normalization.

Installation

First, add RubyLLM to your Gemfile:

gem 'ruby_llm'

Then configure it with your preferred provider (OpenAI, Anthropic, or Google):

# config/initializers/ruby_llm.rb
RubyLLM.setup do |config|
  config.openai_api_key = ENV['OPENAI_API_KEY']
  # or
  config.anthropic_api_key = ENV['ANTHROPIC_API_KEY']
  # or
  config.google_api_key = ENV['GOOGLE_API_KEY']
end

Automatic Content Categorization

Imagine you're building an app that imports blog posts from various sources. Instead of manually tagging them or building complex regex patterns, let the LLM do the heavy lifting:

class PostCategorizer
  CATEGORIES = ['Technology', 'Business', 'Health', 'Entertainment', 'Sports']

  def self.categorize(post)
    prompt = <<~PROMPT
      Categorize this blog post into one of these categories: #{CATEGORIES.join(', ')}

      Title: #{post.title}
      Content: #{post.content.truncate(500)}

      Respond with ONLY the category name, nothing else.
      Format should be "category_name"
      Only 1 category at a time
    PROMPT

    client = RubyLLM.chat(provider: :openai, model: 'gpt-4o-mini')
    response = client.ask(prompt)

    category = response.content

    # You can also have the LLM returning more than one category, I did that for RecAtlas
    # And asked in the prompt for 'Return the response in this format: ["category 1", "category 2"]'
    CATEGORIES.include?(category) ? category : 'Uncategorized'
  end
end

# Usage
post = Post.find(123)
post.update(category: PostCategorizer.categorize(post))

Intelligent Search with Semantic Understanding

Categorizing information is easy, but did you know you could also use the LLM to execute searches against your records? This will be a future feature in RecAtlas, I still need to run a few more tests on that, but in case you're wondering how to implement a search using the LLM for a more "human readable" search format, here's how you can approach that:

PS 1.: This is a suggestion, you should always check the code yourself. The rubyLLM gem can be upgraded and deprecate some of the methods, keep that in mind. PS 2.: This is an example, lines could be reduced, and simplifications can be done, I'm just trying to give you a way to use the LLM.

class SemanticSearcher
  def self.find_relevant_products(user_query)
    # Ask the LLM to translate natural language into structured search parameters
    search_params = extract_search_criteria(user_query)

    # Build the ActiveRecord query using the extracted parameters
    build_product_query(search_params)
  end

  private

  def self.extract_search_criteria(user_query)
    # Define the schema that the LLM should understand
    prompt = <<~PROMPT
      Convert this natural language product search into structured search criteria: "#{user_query}"

      Available product fields:
      - brand (string): Product brand name (e.g., "Nike", "Adidas")
      - color (string): Product color (e.g., "blue", "red", "black")
      - size (string): Product size (e.g., "10", "M", "Large")
      - category (string): Product type (e.g., "shoes", "shirt", "pants")
      - min_price (number): Minimum price filter
      - max_price (number): Maximum price filter
      - name_keywords (string): Keywords to search in product name
      - description_keywords (string): Keywords to search in description

      Return ONLY valid JSON with the extracted criteria. Use null for missing fields.
      Example response format:
      {
        "brand": "Nike",
        "color": "blue",
        "size": "10",
        "category": "shoes",
        "min_price": null,
        "max_price": null,
        "name_keywords": null,
        "description_keywords": null
      }
    PROMPT

    client = RubyLLM.chat(provider: :anthropic, model: 'claude-3-5-haiku-latest')
    response = client.ask(prompt)

    # Parse the JSON response safely
    JSON.parse(response.content)
  rescue JSON::ParserError => e
    Rails.logger.error("Failed to parse LLM search response: #{e.message}")
    {}
  end

  def self.build_product_query(params)
    query = Product.all

    # Apply filters using parameterized queries (SQL injection safe)
    query = query.where("brand ILIKE ?", "%#{params['brand']}%") if params['brand'].present?
    query = query.where("color ILIKE ?", "%#{params['color']}%") if params['color'].present?
    query = query.where("size = ?", params['size']) if params['size'].present?
    query = query.where("category ILIKE ?", "%#{params['category']}%") if params['category'].present?

    # Price range filters
    query = query.where("price >= ?", params['min_price']) if params['min_price'].present?
    query = query.where("price <= ?", params['max_price']) if params['max_price'].present?

    # Keyword search in name or description
    if params['name_keywords'].present?
      query = query.where("name ILIKE ?", "%#{params['name_keywords']}%")
    end

    if params['description_keywords'].present?
      query = query.where("description ILIKE ?", "%#{params['description_keywords']}%")
    end

    query
  end
end

# Usage examples
results = SemanticSearcher.find_relevant_products("Nike shoes, color blue, size 10 US")
# Translates to: Product.where("brand ILIKE ?", "%Nike%").where("color ILIKE ?", "%blue%").where("size = ?", "10")

results = SemanticSearcher.find_relevant_products("red adidas running shoes under $100")
# Translates to: Product.where("brand ILIKE ?", "%adidas%").where("color ILIKE ?", "%red%")
#                        .where("category ILIKE ?", "%shoes%").where("price <= ?", 100)

results = SemanticSearcher.find_relevant_products("looking for a warm winter jacket, prefer black, budget around 50-150 dollars")
# Translates to: Product.where("category ILIKE ?", "%jacket%").where("color ILIKE ?", "%black%")
#                        .where("price >= ?", 50).where("price <= ?", 150)

Data Normalization and Extraction

Dealing with messy user input or third-party data? LLMs excel at extracting structured information from unstructured text:

class AddressNormalizer
  def self.normalize(raw_address)
    prompt = <<~PROMPT
      Extract address components from this text: "#{raw_address}"

      Return ONLY valid JSON with keys: street, city, state, zip_code
      If any component is missing, use null.
    PROMPT

    client = RubyLLM.chat(provider: :anthropic, model: 'claude-3-5-haiku-latest')
    response = client.ask(prompt)

    response.content
  rescue JSON::ParserError
    { error: 'Could not parse address' }
  end
end

# Usage
AddressNormalizer.normalize("123 Main St, Springfield IL 62701")
# => { "street" => "123 Main St", "city" => "Springfield", "state" => "IL", "zip_code" => "62701" }

AddressNormalizer.normalize("lives somewhere on main street in boston")
# => { "street" => "Main Street", "city" => "Boston", "state" => null, "zip_code" => null }

Other Practical Use Cases

Beyond these examples, consider using RubyLLM for:

  • Content moderation: Automatically flag inappropriate user-generated content
  • Sentiment analysis: Classify customer feedback as positive/negative/neutral
  • Email routing: Automatically assign support tickets to the right department
  • Data validation: Check if user input makes logical sense before processing
  • Summary generation: Create concise summaries of long documents or threads
  • Translation and localization: Translate content while maintaining context and tone

The key insight here is that LLMs work best when you give them focused, specific tasks with clear outputs. You don't need to build a full conversational AI to get value from these models - sometimes a simple categorization or extraction job is exactly what your app needs.

Pro tip: For production use, always validate LLM outputs, set reasonable timeouts, and consider caching results for repeated queries. And definitely use the smaller, faster models (like GPT-4o-mini or Gemini Flash) for these focused tasks - they're cheaper and often just as accurate as the flagship models for structured outputs.