# model architecture

Published articles for model architecture.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license

DevFeed: [IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license](<https://devfeed.tech/articles/ibm-releases-sota-granite-time-series-patchtst-fm-r2-model-with-commercial-friendly-license-7266.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ibm-research/ibm-releases-sota-granite-time-series>)

Author: Roman Vaculin; Wesley M Gifford; Jiri Navratil; Chandra Reddy; Ayhan Sebin

Published: 2026-09-09T15:36:24Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [Temporal data](<https://devfeed.tech/topics/temporal-data.md>), [releases](<https://devfeed.tech/topics/releases.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Open Source Models & Datasets](<https://devfeed.tech/topics/open-source-models-datasets.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [apache](<https://devfeed.tech/tags/apache.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [inference](<https://devfeed.tech/tags/inference.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [releases](<https://devfeed.tech/tags/releases.md>), [time-series](<https://devfeed.tech/tags/time-series.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

IBM released Granite Time Series PatchTST-FM-r2, a roughly 385M-parameter time-series foundation model for zero-shot forecasting. The article covers its architecture, probabilistic forecasting, missing-value imputation, benchmark results, licensing, and available reproducibility resources.

### Source excerpt

Time-series foundation models are changing the way forecasting systems are built. Instead of training and maintaining a separate model for every dataset, users can use a pretrained model and generate forecasts zero-shot. IBM has released Granite Time Series PatchTST-FM-r2, the latest model in the Granite TSFM family (github, blog).

## 🗓 This Week In AI Research (1-7 August 26)

DevFeed: [🗓 This Week In AI Research (1-7 August 26)](<https://devfeed.tech/articles/this-week-in-ai-research-1-7-august-26-18282.md>)

Original publisher: [Read original article](<https://www.intoai.pub/p/this-week-in-ai-research-1-7-august>)

Author: Dr. Ashish Bamania

Published: 2026-08-13T19:29:25Z

Content type: article

Language: en

Sources: [Into AI](<https://devfeed.tech/sources/into-ai.md>)

Topics: [AI Research](<https://devfeed.tech/topics/ai-research.md>), [releases](<https://devfeed.tech/topics/releases.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [qwen](<https://devfeed.tech/topics/qwen.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [algorithm](<https://devfeed.tech/tags/algorithm.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [releases](<https://devfeed.tech/tags/releases.md>)

### AI overview

A weekly roundup of AI research and model releases. It highlights Pathway, Bielik AI, and NYU's BDH-CQ reasoning model, which uses in-context learning with recurrent memory and latent-state reasoning, reports ARC-AGI-1 cost-efficiency results, and describes Alibaba's Qwen3.8-Max release and the U-OPSD self-distillation algorithm.

### Source excerpt

The top 10 AI research papers and releases that you must know about this week.

## AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips

DevFeed: [AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips](<https://devfeed.tech/articles/aws-trainium-frontier-competition-co-design-models-and-kernels-on-purpose-built-ai-chips-7612.md>)

Original publisher: [Read original article](<https://www.amazon.science/news/aws-trainium-frontier-competition-co-design-models-and-kernels-on-purpose-built-ai-chips>)

Author: Louise Ping; John Gray; Emily Webber; Josh Longenecker

Published: 2026-08-10T20:23:04Z

Content type: article

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [aws-trainium](<https://devfeed.tech/tags/aws-trainium.md>), [chip-design](<https://devfeed.tech/tags/chip-design.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [kernels](<https://devfeed.tech/tags/kernels.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [neurips](<https://devfeed.tech/tags/neurips.md>), [performance](<https://devfeed.tech/tags/performance.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

AWS Trainium Frontier is a competition for training language models from scratch on Trainium while co-designing architectures, optimizers, training loops, and optional custom kernels under fixed compute and time budgets.

### Source excerpt

A competition with a finalist ceremony during NeurIPS 2026, challenging researchers to train language models from scratch on Trainium, exploring what optimal architectures look like when the hardware changes.

## Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference

DevFeed: [Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference](<https://devfeed.tech/articles/co-designing-ai-model-attention-for-fast-interactive-long-context-inference-6779.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/co-designing-ai-model-attention-for-fast-interactive-long-context-inference/>)

Author: Tanya Lenz

Published: 2026-07-31T22:16:17Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [GPU optimization](<https://devfeed.tech/topics/gpu-optimization.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [memory](<https://devfeed.tech/tags/memory.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [performance](<https://devfeed.tech/tags/performance.md>)

### AI overview

This article examines how co-designing dense attention with GPU execution can improve throughput and interactivity for long-context inference. It analyzes group size, head dimension, sequence length, and the different compute and memory behavior of prefill and decode, including the effects of speculative decoding and prefix caching.

### Source excerpt

As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because...

## Making LLMs faster without sacrificing accuracy

DevFeed: [Making LLMs faster without sacrificing accuracy](<https://devfeed.tech/articles/making-llms-faster-without-sacrificing-accuracy-7603.md>)

Original publisher: [Read original article](<https://www.amazon.science/blog/making-llms-faster-without-sacrificing-accuracy>)

Author: Tao Yu; Youngsuk Park

Published: 2026-05-15T13:00:00Z

Content type: article

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [Language models](<https://devfeed.tech/topics/language-models.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [web applications](<https://devfeed.tech/topics/web-applications.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [chinchilla-scaling-law](<https://devfeed.tech/tags/chinchilla-scaling-law.md>), [conversational-ai](<https://devfeed.tech/tags/conversational-ai.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [grouped-query-attention](<https://devfeed.tech/tags/grouped-query-attention.md>), [hyperparameter-optimization](<https://devfeed.tech/tags/hyperparameter-optimization.md>), [iclr-2026](<https://devfeed.tech/tags/iclr-2026.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-efficiency](<https://devfeed.tech/tags/inference-efficiency.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [llm-optimization](<https://devfeed.tech/tags/llm-optimization.md>), [llms](<https://devfeed.tech/tags/llms.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [network-architectures](<https://devfeed.tech/tags/network-architectures.md>), [scaling-laws](<https://devfeed.tech/tags/scaling-laws.md>), [training](<https://devfeed.tech/tags/training.md>), [transformer-architecture](<https://devfeed.tech/tags/transformer-architecture.md>)

### AI overview

The article presents scaling laws that connect LLM architectural choices to the tradeoff between accuracy and efficiency. It describes how these choices can improve inference throughput without reducing accuracy.

### Source excerpt

A new scaling law that relates particular architectural choices to loss helps identify models that improve throughput by up to 47% with no loss of accuracy.

## Inside Praktika's conversational approach to language learning

DevFeed: [Inside Praktika's conversational approach to language learning](<https://devfeed.tech/articles/inside-praktika-s-conversational-approach-to-language-learning-6613.md>)

Original publisher: [Read original article](<https://openai.com/index/praktika>)

Published: 2026-01-22T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [App](<https://devfeed.tech/topics/app.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [app](<https://devfeed.tech/tags/app.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [building](<https://devfeed.tech/tags/building.md>), [data](<https://devfeed.tech/tags/data.md>), [education](<https://devfeed.tech/tags/education.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [model](<https://devfeed.tech/tags/model.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [product](<https://devfeed.tech/tags/product.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [startup](<https://devfeed.tech/tags/startup.md>)

### AI overview

Praktika uses GPT-powered personalized AI tutors and a multi-agent system to help learners build real-world language fluency through adaptive conversations, progress tracking, and long-term lesson planning.

### Source excerpt

How Praktika uses GPT-4.1 and GPT-5.2 to build adaptive AI tutors that personalize lessons, track progress, and help learners achieve real-world language fluency

## Transformers v5: Simple model definitions powering the AI ecosystem

DevFeed: [Transformers v5: Simple model definitions powering the AI ecosystem](<https://devfeed.tech/articles/transformers-v5-simple-model-definitions-powering-the-ai-ecosystem-7537.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/transformers-v5>)

Author: Lysandre; Arthur Zucker; Cyril Vallez; Vaibhav Srivastav

Published: 2025-12-01T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Transformers](<https://devfeed.tech/topics/transformers.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [community](<https://devfeed.tech/tags/community.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [python](<https://devfeed.tech/tags/python.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [v5](<https://devfeed.tech/tags/v5.md>)

### AI overview

Hugging Face announces Transformers v5, focusing on simpler model definitions and improvements for training, inference, and production. The article describes the library's growth to more than 400 model architectures and over 1.2 billion installs, along with modular design intended to improve maintenance, integration, and collaboration.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## WeatherNext 2: Our most advanced weather forecasting model

DevFeed: [WeatherNext 2: Our most advanced weather forecasting model](<https://devfeed.tech/articles/weathernext-2-our-most-advanced-weather-forecasting-model-6258.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/weathernext-2-our-most-advanced-weather-forecasting-model/>)

Author: The WeatherNext team

Published: 2025-11-17T15:09:23Z

Content type: article

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [google](<https://devfeed.tech/tags/google.md>), [inference](<https://devfeed.tech/tags/inference.md>), [model](<https://devfeed.tech/tags/model.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [none](<https://devfeed.tech/tags/none.md>), [tpu](<https://devfeed.tech/tags/tpu.md>), [vertex-ai](<https://devfeed.tech/tags/vertex-ai.md>)

### AI overview

Google DeepMind and Google Research introduce WeatherNext 2, an AI weather forecasting model that generates hundreds of possible scenarios faster, at higher resolution, and with improved accuracy. Its forecast data is available through Earth Engine and BigQuery, with custom model inference offered through Vertex AI, and its technology is being integrated into several Google products.

### Source excerpt

The new AI model delivers more efficient, more accurate and higher-resolution global weather predictions.

## NVIDIA Nemotron Nano v2: Long-Context Reasoning and Efficient Inference in Smaller Models

DevFeed: [NVIDIA Nemotron Nano v2: Long-Context Reasoning and Efficient Inference in Smaller Models](<https://devfeed.tech/articles/the-future-of-agentic-ai-is-small-35023.md>)

Original publisher: [Read original article](<https://read.theaimerge.com/p/the-future-of-agentic-ai-is-small>)

Author: Alex Razvant

Published: 2025-11-15T14:47:33Z

Content type: article

Language: en

Sources: [Neural Bits](<https://devfeed.tech/sources/neural-bits.md>)

Topics: [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [inference](<https://devfeed.tech/tags/inference.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [moe](<https://devfeed.tech/tags/moe.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>)

### AI overview

This technical article examines NVIDIA's Nemotron Nano v2 Transformer-Hybrid models, focusing on their long-context reasoning, fast inference, and reported performance against larger open models. It also describes the broader Nemotron family's open models, datasets, and fine-tuning recipes for agentic AI systems.

### Source excerpt

How NVIDIA's Nemotron Nano V2 SLM, built for long-context reasoning, fast inference, can compete with models 4-5x its size.

## Fast LoRA inference for Flux with Diffusers and PEFT

DevFeed: [Fast LoRA inference for Flux with Diffusers and PEFT](<https://devfeed.tech/articles/fast-lora-inference-for-flux-with-diffusers-and-peft-7344.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/lora-fast>)

Author: Sayak Paul; Benjamin Bossan

Published: 2025-07-23T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [flux](<https://devfeed.tech/topics/flux.md>), [lora](<https://devfeed.tech/topics/lora.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Hacktoberfest](<https://devfeed.tech/topics/hacktoberfest.md>), [diffusers](<https://devfeed.tech/topics/diffusers.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [peft](<https://devfeed.tech/topics/peft.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [text-to-image](<https://devfeed.tech/topics/text-to-image.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [amd](<https://devfeed.tech/tags/amd.md>), [diffusers](<https://devfeed.tech/tags/diffusers.md>), [diffusion](<https://devfeed.tech/tags/diffusion.md>), [flux](<https://devfeed.tech/tags/flux.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [guide](<https://devfeed.tech/tags/guide.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [lora](<https://devfeed.tech/tags/lora.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [peft](<https://devfeed.tech/tags/peft.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [text-to-image](<https://devfeed.tech/tags/text-to-image.md>)

### AI overview

This article presents an optimization recipe for faster LoRA inference with the Flux.1-Dev text-to-image model using Diffusers and PEFT. The approach addresses LoRA hotswapping and recompilation issues with Flash Attention 3, FP8 quantization from TorchAO, and hotswapping-ready compilation, achieving about a 2.3x speedup while balancing inference speed and memory use.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Mastering Long Contexts in LLMs with KVPress

DevFeed: [Mastering Long Contexts in LLMs with KVPress](<https://devfeed.tech/articles/mastering-long-contexts-in-llms-with-kvpress-7383.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/kvpress>)

Author: Simon Jegou; Maximilian Jeblick

Published: 2025-01-23T08:03:03Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [LLM Techniques](<https://devfeed.tech/topics/llm-techniques.md>), [text-generation](<https://devfeed.tech/topics/text-generation.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [cache](<https://devfeed.tech/tags/cache.md>), [compression](<https://devfeed.tech/tags/compression.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [llms](<https://devfeed.tech/tags/llms.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [model](<https://devfeed.tech/tags/model.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [models](<https://devfeed.tech/tags/models.md>), [text-generation](<https://devfeed.tech/tags/text-generation.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

This article introduces KVPress, an NVIDIA toolkit that applies KV cache compression techniques to make long-context Large Language Models (LLMs) more memory-efficient. It explains how context windows enable in-context retrieval, learning, and extended reasoning, and why KV Cache memory usage grows with context length. The article also describes how KV Cache reuses attention-layer keys and values during autoregressive text generation.

### Source excerpt

TL;DR: KVPress packs the latest KV cache compression techniques, enabling memory-efficient long-context LLMs. 🚀 One of the key features of Large Language Models (LLMs) is their context window--the maximum number of tokens they can process in a single request. As LLMs evolve, their context windows are becoming increasingly larger. Larger context windows unlock incredible possibilities: - In-context retrieval: Seamlessly referencing large amounts of text within a single query.

## Client Time Series Model: a Multi-Target Recommender System based on Temporally-Masked Encoders

DevFeed: [Client Time Series Model: a Multi-Target Recommender System based on Temporally-Masked Encoders](<https://devfeed.tech/articles/client-time-series-model-a-multi-target-recommender-system-based-on-temporally-masked-encoders-29339.md>)

Original publisher: [Read original article](<https://multithreaded.stitchfix.com/blog/2022/10/14/client-time-series-model/>)

Published: 2022-10-14T06:00:00Z

Content type: article

Language: en

Sources: [Stitch Fix](<https://devfeed.tech/sources/stitch-fix.md>)

Topics: [client](<https://devfeed.tech/topics/client.md>), [Time Series](<https://devfeed.tech/topics/time-series.md>), [recommendations](<https://devfeed.tech/topics/recommendations.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [model architecture](<https://devfeed.tech/topics/model-architecture.md>), [maintenance](<https://devfeed.tech/topics/maintenance.md>), [systems](<https://devfeed.tech/topics/systems.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [maintenance](<https://devfeed.tech/tags/maintenance.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [recommendations](<https://devfeed.tech/tags/recommendations.md>), [systems](<https://devfeed.tech/tags/systems.md>), [time-series](<https://devfeed.tech/tags/time-series.md>)

### AI overview

Stitch Fix describes its Client Time Series Model, a sequence-based recommender that estimates the probability of client-item purchases. The model uses a unified client embedding and incorporates the time dimension of client interactions to reduce duplicated models, improve maintainability, and share learning across business lines, regions, and channels.

### Source excerpt

Introduction The foundation of our recommendation stack is a scoring model we call p(sale), which estimates the probability that any given client will purchase any given item. This model has gone through many iterations over the years, from a mixed effects model, to a matrix factorization model, and now to a novel sequence-based model. Internally we call this the Client Time Series Model (aka CTSM) because of its focus on the time-domain of client interactions. This post details our new model, which is a significant improvement for both the quality of our recommendations and the maintainability of our systems. Motivation Before setting out to develop our new model, it was clear that the evolution of our business necessitated a change to our modeling approach. First, the growing variety of recommendations we serve led to an explosion in the number of models we needed to maintain. Each time the business expanded, such as adding Mens, or serving the UK, or adding direct shopping with Freestyle, we responded by forking a new model to serve the new channel. This was necessary because a single domain-agnostic model could not serve the new channels as well as tailored models, but over time it has increased our maintenance burden and cost of iteration. In addition to the system complexity, we also knew we had an opportunity to make better use of important signals. With data and models separated by business line, region, and channel, we had a limited ability to leverage learning across these boundaries. With a unified model, we can more seamlessly use data from US clients to improve recommendations for UK clients, or data from Fixes to improve recommendations in Freestyle. Finally, our previous approaches modeled clients via tabular data. Although they are trained on purchase events that take place in the context of a particular point in time, they did not explicitly consider the time dimension in their understanding of the client's interactions. We believed that there was s

## Transformers for software engineers

DevFeed: [Transformers for software engineers](<https://devfeed.tech/articles/transformers-for-software-engineers-21973.md>)

Original publisher: [Read original article](<https://blog.nelhage.com/post/transformers-for-software-engineers/>)

Author: Nelson Elhage

Published: 2022-04-01T20:00:00Z

Content type: tutorial

Language: en

Sources: [Nelson Elhage](<https://devfeed.tech/sources/nelson-elhage.md>)

Topics: [Transformer architecture](<https://devfeed.tech/topics/transformer-architecture.md>), [Programming](<https://devfeed.tech/topics/programming.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [Reverse Engineering](<https://devfeed.tech/topics/reverse-engineering.md>), [Transformers](<https://devfeed.tech/topics/transformers.md>)

Tags: [anthropic](<https://devfeed.tech/tags/anthropic.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [codex](<https://devfeed.tech/tags/codex.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [gpt-3](<https://devfeed.tech/tags/gpt-3.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [programming](<https://devfeed.tech/tags/programming.md>), [reverse-engineering](<https://devfeed.tech/tags/reverse-engineering.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [storm](<https://devfeed.tech/tags/storm.md>), [transformer-architecture](<https://devfeed.tech/tags/transformer-architecture.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

This tutorial explains Transformer architecture for software engineers, using software engineering and programming perspectives to discuss how GPT-style Transformer models work. It also connects the architecture to interpretability and reverse-engineering efforts.

### Source excerpt

Ever since its introduction in the 2017 paper, Attention is All You Need, the Transformer model architecture has taken the deep-learning world by storm. Initially introduced for machine translation, it has become the tool of choice for a wide range of domains, including text, audio, video, and others. Transformers have also driven most of the massive increases in model scale and capability in the last few years. OpenAI's GPT-3 and Codex models are Transformers, as are DeepMind's Gopher models and many others.

## Deep NLP-based Recommenders at Finn.no

DevFeed: [Deep NLP-based Recommenders at Finn.no](<https://devfeed.tech/articles/deep-nlp-based-recommenders-at-finn-no-32019.md>)

Original publisher: [Read original article](<https://tech.finn.no2017/09/08/NLP-based-recommenders-at-finn/>)

Author: Simen Eide

Published: 2017-09-08T13:56:49Z

Content type: tutorial

Language: en

Sources: [Finn.no](<https://devfeed.tech/sources/finn-no.md>)

Topics: [Natural language processing](<https://devfeed.tech/topics/nlp.md>), [Deep learning](<https://devfeed.tech/topics/deep-learning.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [recommendations](<https://devfeed.tech/topics/recommendations.md>), [Keras](<https://devfeed.tech/topics/keras.md>), [Hackathon](<https://devfeed.tech/topics/hackathon.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [hackathon](<https://devfeed.tech/tags/hackathon.md>), [keras](<https://devfeed.tech/tags/keras.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [neural-networks](<https://devfeed.tech/tags/neural-networks.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [recommendations](<https://devfeed.tech/tags/recommendations.md>)

### AI overview

The article describes a FINN.no hackathon project exploring deep NLP-based recommendations for classified ads. The team used categorized ad data, word embeddings, and a convolutional neural network architecture to model similarity, but the supplied text does not include the final performance results.

### Source excerpt

During a hackathon at FINN.no, we figured we wanted to learn more about deep NLP-models. FINN.no has a large database with ads of people trying to sell stuff (around 1 million active ads at any time), and they are categorized into a category tree with three or four layers. For example, full suspension bikes can be found under "Sport and outdoor activities" / "Bike sport" / "Full suspension bikes". In our daily jobs we are working on recommendations. There, we already have a content based (tf-idf) recommender build on Solr's More Like This. It seems to work well in areas where our collaborative filtering approaches does not. Would it be possible to build a deep learning NLP-model of similar performance? To achieve a measure of similarity, building a classifier of the previously mentioned categories seemed like a good choice, since we already had a lot of pre-existing data. The NLP team at Schibsted had already tokenized around six million ads as well as trained a word2vec model for us - we were ready to roll! Some preprocessing still had to be done. We ran through all ads, concatenated the title and description strings, and after a quick look at the data took the first 15 words of each ad. Model architecture proposed by the paper Our initial experiments were done with a simple "Bag of words" model included in the Keras repository, but we promptly switched over to "Convolutional Neural Networks for Sentence Classification" based architecture after hearing about it from our colleague, Tobias. By looking at the first 15 words of the ad, and using 200 dimensional embeddings for each word, our input is transformed into a 15x200 matrix. We apply three different convolutions on each document. The three convolutions looks at 2, 3 and 4 words (kernel sizes) in each convolution. It then max-pools each over the whole document, so that you end up with one value per document per convolution. For each kernel size you do 100 different filters. Finally you add a dense layer for clas