Recently, I have been investing quite some time building some features for my personal project RecAtlas, and I had one challenge in my hands regarding the standardization of the categories of the activities. At first, I was thinking about implementing some FuzzySearch with PostgreSQL and create some pre-defined categories, and then, a data migration, and a categorization class for the new activities fetched from the API, but then, I realized: "Why not use one of the things that AI does best, which is categorization of information?".
When people think about integrating LLMs into their Rails apps, they usually jump straight to building chat interfaces. But here's the thing - some of the most powerful applications of LLMs are way more subtle. Let's talk about using RubyLLM to solve everyday problems like categorization, intelligent search, and data normalization.
Installation
First, add RubyLLM to your Gemfile:
gem 'ruby_llm'
Then configure it with your preferred provider (OpenAI, Anthropic, or Google):
# config/initializers/ruby_llm.rb
RubyLLM.setup do |config|
config.openai_api_key = ENV['OPENAI_API_KEY']
# or
config.anthropic_api_key = ENV['ANTHROPIC_API_KEY']
# or
config.google_api_key = ENV['GOOGLE_API_KEY']
end
Automatic Content Categorization
Imagine you're building an app that imports blog posts from various sources. Instead of manually tagging them or building complex regex patterns, let the LLM do the heavy lifting:
class PostCategorizer
CATEGORIES = ['Technology', 'Business', 'Health', 'Entertainment', 'Sports']
def self.categorize(post)
prompt = <<~PROMPT
Categorize this blog post into one of these categories: #{CATEGORIES.join(', ')}
Title: #{post.title}
Content: #{post.content.truncate(500)}
Respond with ONLY the category name, nothing else.
Format should be "category_name"
Only 1 category at a time
PROMPT
client = RubyLLM.chat(provider: :openai, model: 'gpt-4o-mini')
response = client.ask(prompt)
category = response.content
# You can also have the LLM returning more than one category, I did that for RecAtlas
# And asked in the prompt for 'Return the response in this format: ["category 1", "category 2"]'
CATEGORIES.include?(category) ? category : 'Uncategorized'
end
end
# Usage
post = Post.find(123)
post.update(category: PostCategorizer.categorize(post))
Intelligent Search with Semantic Understanding
Categorizing information is easy, but did you know you could also use the LLM to execute searches against your records? This will be a future feature in RecAtlas, I still need to run a few more tests on that, but in case you're wondering how to implement a search using the LLM for a more "human readable" search format, here's how you can approach that:
PS 1.: This is a suggestion, you should always check the code yourself. The rubyLLM gem can be upgraded and deprecate some of the methods, keep that in mind. PS 2.: This is an example, lines could be reduced, and simplifications can be done, I'm just trying to give you a way to use the LLM.
class SemanticSearcher
def self.find_relevant_products(user_query)
# Ask the LLM to translate natural language into structured search parameters
search_params = extract_search_criteria(user_query)
# Build the ActiveRecord query using the extracted parameters
build_product_query(search_params)
end
private
def self.extract_search_criteria(user_query)
# Define the schema that the LLM should understand
prompt = <<~PROMPT
Convert this natural language product search into structured search criteria: "#{user_query}"
Available product fields:
- brand (string): Product brand name (e.g., "Nike", "Adidas")
- color (string): Product color (e.g., "blue", "red", "black")
- size (string): Product size (e.g., "10", "M", "Large")
- category (string): Product type (e.g., "shoes", "shirt", "pants")
- min_price (number): Minimum price filter
- max_price (number): Maximum price filter
- name_keywords (string): Keywords to search in product name
- description_keywords (string): Keywords to search in description
Return ONLY valid JSON with the extracted criteria. Use null for missing fields.
Example response format:
{
"brand": "Nike",
"color": "blue",
"size": "10",
"category": "shoes",
"min_price": null,
"max_price": null,
"name_keywords": null,
"description_keywords": null
}
PROMPT
client = RubyLLM.chat(provider: :anthropic, model: 'claude-3-5-haiku-latest')
response = client.ask(prompt)
# Parse the JSON response safely
JSON.parse(response.content)
rescue JSON::ParserError => e
Rails.logger.error("Failed to parse LLM search response: #{e.message}")
{}
end
def self.build_product_query(params)
query = Product.all
# Apply filters using parameterized queries (SQL injection safe)
query = query.where("brand ILIKE ?", "%#{params['brand']}%") if params['brand'].present?
query = query.where("color ILIKE ?", "%#{params['color']}%") if params['color'].present?
query = query.where("size = ?", params['size']) if params['size'].present?
query = query.where("category ILIKE ?", "%#{params['category']}%") if params['category'].present?
# Price range filters
query = query.where("price >= ?", params['min_price']) if params['min_price'].present?
query = query.where("price <= ?", params['max_price']) if params['max_price'].present?
# Keyword search in name or description
if params['name_keywords'].present?
query = query.where("name ILIKE ?", "%#{params['name_keywords']}%")
end
if params['description_keywords'].present?
query = query.where("description ILIKE ?", "%#{params['description_keywords']}%")
end
query
end
end
# Usage examples
results = SemanticSearcher.find_relevant_products("Nike shoes, color blue, size 10 US")
# Translates to: Product.where("brand ILIKE ?", "%Nike%").where("color ILIKE ?", "%blue%").where("size = ?", "10")
results = SemanticSearcher.find_relevant_products("red adidas running shoes under $100")
# Translates to: Product.where("brand ILIKE ?", "%adidas%").where("color ILIKE ?", "%red%")
# .where("category ILIKE ?", "%shoes%").where("price <= ?", 100)
results = SemanticSearcher.find_relevant_products("looking for a warm winter jacket, prefer black, budget around 50-150 dollars")
# Translates to: Product.where("category ILIKE ?", "%jacket%").where("color ILIKE ?", "%black%")
# .where("price >= ?", 50).where("price <= ?", 150)
Data Normalization and Extraction
Dealing with messy user input or third-party data? LLMs excel at extracting structured information from unstructured text:
class AddressNormalizer
def self.normalize(raw_address)
prompt = <<~PROMPT
Extract address components from this text: "#{raw_address}"
Return ONLY valid JSON with keys: street, city, state, zip_code
If any component is missing, use null.
PROMPT
client = RubyLLM.chat(provider: :anthropic, model: 'claude-3-5-haiku-latest')
response = client.ask(prompt)
response.content
rescue JSON::ParserError
{ error: 'Could not parse address' }
end
end
# Usage
AddressNormalizer.normalize("123 Main St, Springfield IL 62701")
# => { "street" => "123 Main St", "city" => "Springfield", "state" => "IL", "zip_code" => "62701" }
AddressNormalizer.normalize("lives somewhere on main street in boston")
# => { "street" => "Main Street", "city" => "Boston", "state" => null, "zip_code" => null }
Other Practical Use Cases
Beyond these examples, consider using RubyLLM for:
- Content moderation: Automatically flag inappropriate user-generated content
- Sentiment analysis: Classify customer feedback as positive/negative/neutral
- Email routing: Automatically assign support tickets to the right department
- Data validation: Check if user input makes logical sense before processing
- Summary generation: Create concise summaries of long documents or threads
- Translation and localization: Translate content while maintaining context and tone
The key insight here is that LLMs work best when you give them focused, specific tasks with clear outputs. You don't need to build a full conversational AI to get value from these models - sometimes a simple categorization or extraction job is exactly what your app needs.
Pro tip: For production use, always validate LLM outputs, set reasonable timeouts, and consider caching results for repeated queries. And definitely use the smaller, faster models (like GPT-4o-mini or Gemini Flash) for these focused tasks - they're cheaper and often just as accurate as the flagship models for structured outputs.