Invest In Crypto News
  • Home
  • Latest News
    • Bitcoin News
    • Altcoin News
    • Ethereum News
    • Blockchain News
    • Doge News
    • NFT News
    • Video
    • Market Analysis
    • Business
    • Finance
    • Politics
    • Mining
    • Regulation
    • Technology
  • Top 10 Cryptos
  • Market Cap List
  • IC DAO
  • Donations
  • Contact
  • Buy Crypto
No Result
View All Result
Invest In Crypto News
  • Home
  • Latest News
    • Bitcoin News
    • Altcoin News
    • Ethereum News
    • Blockchain News
    • Doge News
    • NFT News
    • Video
    • Market Analysis
    • Business
    • Finance
    • Politics
    • Mining
    • Regulation
    • Technology
  • Top 10 Cryptos
  • Market Cap List
  • IC DAO
  • Donations
  • Contact
  • Buy Crypto
No Result
View All Result
Invest In Crypto News
No Result
View All Result

IBM Research Unveils Cost-Effective AI Inferencing with Speculative Decoding

CryptoExpert by CryptoExpert
June 24, 2024
in Blockchain News
0
IBM Research Unveils Cost-Effective AI Inferencing with Speculative Decoding
  • Facebook
  • Twitter
  • Pinterest


You might also like

ZetaChain Holders Approve L1 Shutdown Plan, Solana Migration

COIN Price Prediction: Squeezed Against the Ceiling — Bull Run or Bull Trap at $192?

HOOD Price Prediction: Bulls Are Right But the Easy Money Is Gone — $130 Is the Real Test






IBM Research has announced a significant breakthrough in AI inferencing, combining speculative decoding with paged attention to enhance the cost performance of large language models (LLMs). This development promises to make customer care chatbots more efficient and cost-effective, according to IBM Research.

In recent years, LLMs have improved the ability of chatbots to understand customer queries and provide accurate responses. However, the high cost and slow speed of serving these models have hindered broader AI adoption. Speculative decoding emerges as an optimization technique to accelerate AI inferencing by generating tokens faster, which can reduce latency by two to three times, thereby improving customer experience.

Despite its advantages, reducing latency traditionally comes with a trade-off: decreased throughput, or the number of users that can simultaneously utilize the model, which increases operational costs. IBM Research has tackled this challenge by cutting the latency of its open-source Granite 20B code model in half while quadrupling its throughput.

Speculative Decoding: Efficiency in Token Generation

LLMs use a transformer architecture, which is inefficient at generating text. Typically, a forward pass is required to process each previously generated token before producing a new one. Speculative decoding modifies this process to evaluate several prospective tokens simultaneously. If these tokens are validated, one forward pass can generate multiple tokens, thus increasing inferencing speed.

Betfury

This technique can be executed by a smaller, more efficient model or part of the main model itself. By processing tokens in parallel, speculative decoding maximizes the efficiency of each GPU, potentially doubling or tripling inferencing speed. Initial introductions of speculative decoding by DeepMind and Google researchers utilized a draft model, while newer methods, such as the Medusa speculator, eliminate the need for a secondary model.

IBM researchers adapted the Medusa speculator by conditioning future tokens on each other rather than on the model’s next predicted token. This approach, combined with an efficient fine-tuning method using small and large batches of text, aligns the speculator’s responses closely with the LLM, significantly boosting inferencing speeds.

Paged Attention: Optimizing Memory Usage

Reducing LLM latency often compromises throughput due to increased GPU memory strain. Dynamic batching can mitigate this but not when speculative decoding is also competing for memory. IBM researchers addressed this by employing paged attention, an optimization technique inspired by virtual memory and paging concepts from operating systems.

Traditional attention algorithms store key-value (KV) sequences in contiguous memory, leading to fragmentation. Paged attention, however, divides these sequences into smaller blocks, or pages, that can be accessed as needed. This method minimizes redundant computation and allows the speculator to generate multiple candidates for each predicted word without duplicating the entire KV-cache, thus freeing up memory.

Future Implications

IBM has integrated speculative decoding and paged attention into its Granite 20B code model. The IBM speculator has been open-sourced on Hugging Face, enabling other developers to adapt these techniques for their LLMs. IBM plans to implement these optimization techniques across all models on its watsonx platform, enhancing enterprise AI applications.

Image source: Shutterstock



Source link

  • Facebook
  • Twitter
  • Pinterest
CryptoExpert

CryptoExpert

Recommended For You

ZetaChain Holders Approve L1 Shutdown Plan, Solana Migration

by CryptoExpert
September 21, 2026
0
Cointelegraph

ZetaChain tokenholders have approved a plan to wind down the project’s layer-1 blockchain and migrate its native ZETA token to Solana. Governance proposal 68 passed Sunday with 99.4% support...

Read more

COIN Price Prediction: Squeezed Against the Ceiling — Bull Run or Bull Trap at $192?

by CryptoExpert
September 21, 2026
0
COIN Price Prediction: Squeezed Against the Ceiling — Bull Run or Bull Trap at $192?

James Ding Sep 20, 2026 13:05 Coinbase Global (COIN) is trading at $192.90, pressed hard against its upper Bollinger Band with MACD momentum exhausted...

Read more

HOOD Price Prediction: Bulls Are Right But the Easy Money Is Gone — $130 Is the Real Test

by CryptoExpert
September 20, 2026
0
HOOD Price Prediction: Bulls Choking Below $100 — Smart Money Is Loaded but the Tape Is Lying

Timothy Morano Sep 20, 2026 13:14 Robinhood (HOOD) is pulling back 2.2% to $117.15 with momentum stalling at a critical inflection zone. With Wall...

Read more

SHIB Price Prediction: Dead-Cat Bounce or Real Recovery? The $0.0000057 Wall Will Decide Everything

by CryptoExpert
September 20, 2026
0
SHIB Price Prediction: Dead-Cat Bounce or Real Recovery? The $0.0000057 Wall Will Decide Everything

Ted Hisokawa Sep 20, 2026 10:12 SHIB is trading at $0.00000539, clinging to its 50-day moving average after a brutal 58% year-on-year wipeout. With...

Read more

PLTR Price Prediction: AIPCon Momentum Meets a $250 Wall Street Target — Can Bulls Clear $180?

by CryptoExpert
September 20, 2026
0
PLTR Price Prediction: Blowout Earnings Meet Overbought Technicals — Brace for a $165–$195 Decision Point

Ted Hisokawa Sep 19, 2026 13:18 Palantir is trading at $176.87 with institutional money firmly long and UBS targeting $250, but momentum has flatlined...

Read more
Next Post
Market Turbulence Continues as Crypto Outflows Reach $1.2 Billion

Market Turbulence Continues as Crypto Outflows Reach $1.2 Billion

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Browse by Category

  • Altcoin News
  • Bitcoin News
  • Blockchain News
  • Business
  • Doge News
  • Ethereum News
  • Finance
  • Market Analysis
  • Mining
  • NFT News
  • Politics
  • Regulation
  • Technology
  • Trending Cryptos
  • Video

Sitemap

  • Market Cap
  • Donations
  • Trading
  • Mining
  • Contact

Legal Information

  • Privacy Policy
  • Anti-Spam Policy
  • Copyright Notice
  • DMCA Compliance
  • Social Media Disclaimer
  • Terms Of Service

Categories

  • Altcoin News
  • Bitcoin News
  • Blockchain News
  • Business
  • Doge News
  • Ethereum News
  • Finance
  • Market Analysis
  • Mining
  • NFT News
  • Politics
  • Regulation
  • Technology
  • Trending Cryptos
  • Video

© Copyright 2024 InvestInCryptoNews.com

No Result
View All Result
  • Home
  • Latest News
    • Bitcoin News
    • Altcoin News
    • Ethereum News
    • Blockchain News
    • Doge News
    • NFT News
    • Video
    • Market Analysis
    • Business
    • Finance
    • Politics
    • Mining
    • Regulation
    • Technology
  • Top 10 Cryptos
  • Market Cap List
  • IC DAO
  • Donations
  • Contact
  • Buy Crypto

© Copyright 2024 InvestInCryptoNews.com

This website is using cookies to improve the user-friendliness. You agree by using the website further.

Privacy policy
bitcoin
Bitcoin (BTC) $ 85,321.00
ethereum
Ethereum (ETH) $ 2,728.28
tether
Tether (USDT) $ 0.999695
bnb
BNB (BNB) $ 789.91
xrp
XRP (XRP) $ 1.49
usd-coin
USDC (USDC) $ 0.999753
solana
Solana (SOL) $ 116.78
tron
TRON (TRX) $ 0.34435
staked-ether
Lido Staked Ether (STETH) $ 2,265.05
zcash
Zcash (ZEC) $ 1,547.46

Pin It on Pinterest

Are you sure want to unlock this post?
Unlock left : 0
Are you sure want to cancel subscription?