分類彙整: GPU

CUDA 2.2 beta with zero-copy access

http://forums.nvidia.com/index.php?showtopic=92112

CUDA vs ATI Stream comparison

http://forums.nvidia.com/index.php?showtopic=92416

CUDA 2.2 features

A brief overview of CUDA 2.2 beta features:

– Zero-copy support (see this thread for more information)

http://forums.nvidia.com/index.php?showtopic=92290

Cuda 2.2 / Zero-copy access

The Cuda 2.2 release notes describe the new zero-copy feature:

* Zero-copy access to pinned system memory

+ Allows MCP7x and GT200 and later GPUs to use system memory withoutcopying to dedicated (video) memory for significant perf improvement.

MCP7x和GT200現在可以直接DMA主記憶體。

G9x看來是不行了。

– Asynchronous memcpy on Vista/Server 2008

– Texturing from pitchlinear memory

– cuda-gdb for 64-bit Linux (it is pretty great)

– OGL interop performance improvements

– CUDA profiler supports a lot more counters on GT200. I think this includes memory bandwidth counters (counters for each transaction size) and instruction count. In other words, you can very easily determine if you’re bandwidth limited or compute limited, which makes it far more useful than it used to be.

– CUDA profiler works on Vista

– >4GB of pinned memory in a single allocation (except in Vista, where the limit is still 256MB per allocation, but I think this is going to be raised between now and the final release)

– Blocking sync for all platforms. Whether this made it into the headers for the beta, I’m not entirely sure–I’ve heard conflicting reports and need to check this afternoon. Basically, it’s a context creation flag where instead of spinlocking or spinlocking+yielding when a thread is waiting for the GPU, the thread will sleep and the driver will wake it up when the event has completed. It’s not the default mode because you’re at the mercy of the OS thread scheduler which will sometimes increase latency, but if you want to minimize CPU utilization, it’s very nice.

– Officially supports Ubuntu 8.10, RHEL 5.3, Fedora 10

——-

http://pc.watch.impress.co.jp/docs/2009/0402/nvidia.htm

NVIDIA、240SP搭載のパフォーマンスGPU「GeForce GTX 275」

~ゲーム性能を向上させた新ドライバも公開

http://pc.watch.impress.co.jp/docs/2009/0402/tawada168.htm

GT200コア採用の新ハイエンドGPU「GeForce GTX 275」

~新機能を実装したGeForce Release 185もリリース

http://heresy.spaces.live.com/blog/cns!E0070FB8ECF9015F!6442.entry?wa=wsignin1.0&sa=517742921

2009/4/3

ForceWare 185.63、Graphics Plus Power Pack #3 & nVidia CUDA 2.2 封測?

http://bb.watch.impress.co.jp/cda/news/25450.html

iPhone版「Skype」アプリ、100万ダウンロードを突破

http://plusd.itmedia.co.jp/pcuser/articles/0904/02/news042.html

GPUの2009シーズンは本日開幕──Radeon HD 4890でベンチを走らせる (1/2)

http://www.ddj.com/architect/216402188

A First Look at the Larrabee New Instructions (LRBni)

http://software.intel.com/file/15542

http://software.intel.com/file/15545

LRBni

http://plusd.itmedia.co.jp/lifestyle/articles/0812/24/news031.html

爆速H.264変換×超解像カード「FIRECODER Blu」を駆る

http://plusd.itmedia.co.jp/pcuser/articles/0812/08/news040.html

大盛況だった「SpursEngine」イベント――SDKの無償配布も告知

GeForce GTX280M發表

http://www.nvidia.com/object/product_geforce_gtx_280m_us.html

GPU Engine Specs:

Processor Cores 128

Memory Interface Width 256-bit

叫做GTX280M卻用G92你們該死XD

G92會不會太有競爭力….

http://www.nvidia.com/object/product_quadro_nvs_450_us.html

Quaro NVS 450

有4個Display Port看起來還蠻猛的。

http://cancoffee2.at.webry.info/200805/article_82.html

General Purpose GPU と CUDA 開発環境

General Purpose GPU,略して GPGPU らしい.

http://www.atmarkit.co.jp/news/200803/06/cuda.html



物理モデル MIDI 音源も,GPU を使うようにすれば,いまよりレスポンスがよくなるのかもー.

あるいは,レスポンスはさておいて,音質があがるとか,リアリティが増す,とか,できるかもー.

ピアノだったら,ペダル踏度に応じて出音が変わるとか.

—-

http://gigazine.net/index.php?/news/comments/20070203_sampleswap/

4.6GBものサウンド素材が無料ダウンロードできる「SampleSwap」

—-

http://www.atmarkit.co.jp/news/200903/18/twitter.html

検索エンジンとしてのTwitter

http://www.atmarkit.co.jp/news/200902/09/weekly.html

Google記事の裏にTwitterあり(@ITNews)

http://www.atmarkit.co.jp/news/200902/13/twitter.html

Twitter、その自信の根拠は(@ITNews)

http://www.atmarkit.co.jp/news/200903/09/eweek.html

Twitterの検索機能はグーグルを脅かすか(@ITNews)

老兵不死的G92:GeForce GTS 250

講好聽一點是這樣,講難聽一點就是看到膩了w

4) GTS 250:128sp(1836MHz)、738/2200MHz、256-bit、55nm

3) 9800 GTX+:128sp(1836MHz)、738/2200MHz、256-bit、55nm

2) 9800 GTX:128sp(1688MHz)、675/2200MHz、256-bit、65nm

1) 8800 GTS 512:128sp(1625MHz)、650/1940MHz、256-bit、65nm

不過話說回來,的確用GT200去改成中階不見得會比G92好….

G92:NVIDIA GeForce GTS 150

G94:NVIDIA GeForce GT 130

G96:NVIDIA GeForce GT 120

—-

http://pc.watch.impress.co.jp/docs/2009/0304/apple01.htm

アップル、NVIDIAベースになったMac mini

http://plusd.itmedia.co.jp/pcuser/articles/0903/03/news118.html

システムを一新してFireWire 800とMini DisplayPortを備えた新型Mac miniが登場

Core2Duo(2GHz、3MB L2、667MHz FSB)和9400M嗎….

雖說不是Ion有點可惜,不過想想9400M本身是一樣的,所以某種意味上傳言其實很接近。

(NVIDIA Ion = Atom + 9400M)

而且即使是Mac Mini,desktop用Atom老實說問題還蠻大的,畢竟Mac mini是省空間desktop不是nettop….

GeForce 9000系改名為GeForce GTS 1xx

http://www.tcmagazine.com/comments.php?shownews=22074&catid=3

Nvidia’s new naming scheme revealed – TechConnect Magazine

– NVIDIA_G92.DEV_0615.1 = “NVIDIA GeForce GTS 150” ← GeForce 9800

– NVIDIA_G94.DEV_0626.1 = “NVIDIA GeForce GT 130” ← GeForce 9600

– NVIDIA_G96.DEV_0646.1 = “NVIDIA GeForce GT 120” ← GeForce 9500

這是反映G8x/G9x = GT1xx這個背後的意義。

但是搞得光G98這一代改名三次(8800GT –> 9800GT –> GeForce GTS150)…..只是讓印象更壞而已。

以史上賣得最快的高階卡(一個月兩百萬張)的8800GT來說,實在是很爛尾風的狀況,讓人不知如何是好。

—-

http://pc.watch.impress.co.jp/docs/2008/1003/ubiq230.htm

ネットブックが外付けGPUを搭載しない理由

對NVIDIA來說CUDA是讓GPU在各個產品階層都可以擴充使用群的一個重大進展。

不過在開花結果之前很難被當成一回事,而且還因此犧牲了繪圖市場的進步速度,導致08年的再逆轉。

CUDA畢竟是不推不行的東西所以沒話說,只能說Intel真的很大_A_,大到不行。

[CEDEC08]驕兵必敗與GT206/GT212

http://pc.watch.impress.co.jp/docs/2008/0912/kaigai467.htm

R600のパフォーマンスを低価格帯にもたらす「ATI Radeon HD 4600」

看了這個就會感覺到差不多有當初Radeon9550的衝擊力了。orz

http://www.4gamer.net/games/032/G003263/20080910062/

[CEDEC 2008#07]NVIDIA開発の鉄人総ざらい,ついにリアルタイムヘアシミュレーションが

這邊開始看到一些努力反擊的感覺….

其實和當初GeForce4的成功造成GeForceFX的失敗相同,GeForce8的大成功似乎也引來了GeForce GTX200系列的失敗。

一邊是想走高programmable性、一邊是想強化CUDA,都是”開始偏副業”的時後被追上;

話說有趣的事情是這兩次都是改數字的時候出槌。XD

http://www.firingsquad.tw/fstw/news/tonews.action?n=7635

總之目前看來GT200的兩個後繼品,GT206是55nm、GT212是40nm,替代關係是GT260->GT206,GT280->GT212。

然後9800GTX+後面還會補上”HDV”的功能,指的應該是55nm G98加入的PureVideo3(主要是VC-1支援)。

—-

話說因為這回RV770/R700的大勝利,很多人開始追捧單卡多晶片是要繼續達成性能成長必然的作法….

不過我個人其實是比較支持單卡單晶片、非必要(多卡)不使用多晶片的做法….效率畢竟還是會有落差。

這回AMD如果真的確定要賣fab,那麼先前想的一堆AMD GPU透過IBM取得的製程優勢可能性就又開始要檢討了。

反過來說,他們積極衝multi-chip也許也與這點有關係?

100M ray/s 的CUDA ray tracer?!

http://forum.beyond3d.com/showthread.php?t=49689

某位仁兄在B3D丟自己寫的ray tracer性能數據。

http://bouliiii.blogspot.com/

http://bouliiii.blogspot.com/2008/08/real-time-ray-tracing-with-cuda-100.html

With the last optimizations I made, I think that between 12 millions rays / s and 40 millions rays/s may be computed on a GeForce 8800GT on the demo I give in the code.

作者來頭:

http://www710.univ-lyon1.fr/~bsegovia/

demo:

http://www710.univ-lyon1.fr/~bsegovia/demos/radius-cuda.zip

哇塞,8800GT跑12M~40M ray/s?那GTX280不就有機會破100M ray/s?

[EDIT]

把Larrabee的數字搞錯啦~是一個需要「4M rays」的場景….

所以16core跑41.16fps、32core跑71.63fps,應該代表的是164M ray/s和286.52M ray/s。

這樣的話GTX280的數字當然就變得比較合理了。

NVIDIA Force within推出

http://www.nvidia.com/content/forcewithin/us/download.asp

Geforce power pack:download

不過對手是4870X2….

http://pc.watch.impress.co.jp/docs/2008/0812/amd.htm

AMD、2GPU搭載の「ATI Radeon HD 4870 X2」~カード1枚で2.4TFLOPSの演算性能を実現

http://pc.watch.impress.co.jp/docs/2008/0812/tawada149.htm

AMDの新アーキテクチャ版マルチGPUカード 「Radeon HD 4870 X2」

兩顆合起來也才520mm^2,比GT200的576mm^2還小得多。

但是要說單顆與雙晶片卡無法抗衡的話,其實NVIDIA也有個9800GX2在那邊,那為什麼會這麼慘呢?

我覺得重點還是在承載能力就是了….

今天如果拿兩顆G92b做成GX2、一樣放2GB GDDR5在上頭的話,表現會如何呢?

考慮NVIDIA目前推Force Within,他們似乎是考慮對既有GeForce user提出升級號召,這樣只要PhysX使用的title,有任意一張副卡輔助的話效率都很可能會比組成SLI還好,畢竟要一個晶片通吃實在是太勉強了。

可是另外一方面,考慮RV730和RV670規模幾乎一樣,只是記憶體介面變成128bit,晶片本身卻壓到和9500GT差不多的報價,可以看得出來ATI一定有少賺….算是壓縮利潤在硬衝。

如果NVIDIA也全跟的話,GPU市場價格可能會徹底崩盤,那存貨壓力比較嚴重的N系AIC損失會比較大。

此外,高階平台的話,ATI的driver quality似乎還跟不上:

http://enthusiast.hardocp.com/image.html?image=MTIxNTk3NjMwNHdjT0psaWNvM3pfNV8yX2wuZ2lm

http://enthusiast.hardocp.com/image.html?image=MTIxNTk3NjMwNHdjT0psaWNvM3pfNV8zX2wuZ2lm

Crysis 1.2.1(DX10)

1920×1200 8xAA/16AF

Radeon4870X2 CrossFire-X vs GTX280 SLI

Big-Bang2 內容

http://www.geeks3d.com/?p=397

NVIDIA Big Bang II – OpenGL 3.0

ForceWare180開始提供把閒置的2nd GPU設為物理專用….拿舊卡跑多螢幕display的同時也不會完全閒置,還是可以跑GPGPU。

這功能看起來沒有必要透過chipset限制,不過如果NVIDIA這樣搞就爆笑了。

可以做的事情也很明顯:CUDA限制繪圖與GPGPU同時執行的時候,GPGPU本身工作會因為OS的Watchdog的關係,最長只能run 5sec;但是一個程式需要run五秒完全不作任何輸出其實是很怪的,應該是跑一下丟一點結果出來monitor、再跑一下再丟一些的狀況比較好;而如果你完全都要把GPU吃死的話,那的確是準備另外一張比較適合。

目前的話,如果做video encoder的話可能就會面臨這樣的問題:比方說把畫面擷取下來即時壓縮成H.264、扔出去給Media Extender來跑,類似PS3的Remote Play這樣的功能的話,Encoder幾乎是抓固定的framerate,並且一直在run的關係,那就最好是另外準備一張卡。

如果有這種solution的話我是不是可以準備把9600GT拿來當成純compute device了XD

而且事實上,光是9500GT就已經有相當的運算能力….

http://pc.watch.impress.co.jp/docs/2008/0730/nvidia.htm

NVIDIA、ミドルレンジ向けGPU「GeForce 9500 GT」