黨大談:結訓假

雖然只能回來三天,還是會想回家…. _A_

把HA8000的抽屜帶回來、小黑送修、家裡的ADSL、FTP硬碟增設、換Power….其實這幾天大概還是超忙的。

不過,連電視沒聲音、小燈泡燒掉也不管,你們真的把我當水電工嗎。_A_)a

比較慘的是小黑送修,現在Ultrabay抽不出來。=w=
上回底盤裝死硬是不修,這回看起來是逃不過了….

Cell 再修改、倍精度效能提升至單精度的1/2

來自IBM Cell Developer Forum的消息

1. 本月結束為止Cell 所需軟體與函式庫發放
2. IBM再對Cell做了點修改,讓原來只有單精度約1/10~1/14速度的倍精度浮點,加快到單精度的1/2。

IBM那些怪物又開始惡搞了…..
這下waterball兄先前在巴哈PS3版下的結論要準備收回了,雖然他自己顯然很高興。

不過到現在還有空慢慢改CPU細節設計,我看他們真的沒打算明年春天讓PS3上市了….

啊,我還是覺得只有一顆Cell會被製造出來,PS3拿到的Cell 不會和IBM/Toshiba到時候採用的Cell有什麼不同,所以這個DP效能改進我相信PS3會受惠。

比較應該注意的,是未完成的不只是Cell,還有RSX….

6800GS-保留實力至今的NV42

為了對抗X1600XT,NVIDIA宣佈會在短期內推出6800GS….
以NV42為基礎的產品,定價USD 249,等同X1600XT。

6800GS規格上目前不明,不過原則上這邊猜測應該與FX3450雷同:
425/500、256MB GDDR3。

總之,NV42與RV530之間的戰鬥應該精采可期,這也是RV530證明自己架構優越的絕佳機會:
425/500 vs 590/500,記憶體完全相同。
NV42-5vs12ps、12ROP、12TMU、256bit (4x64bit)、202M 電晶體。
RV530-5vs12ps、4ROP、4TMU、128bit (4x32bit + Ring bus)、157M。
Ultra Threading + Ring Bus Memory Controller + 時脈優勢,能不能打贏 TMU差距 & 記憶體頻寬差距呢?

這是越級挑戰啊。XD

話說"技術上的優勢",ATI現在比較樂於公開討論自己的優勢,這是好進展:
如B3D開始有ATI的人員公開出來討論"Toy shop"這個demo。

http://www.beyond3d.com/forum/showpost.php?p=593755&postcount=78

NV4x/G7x無法順利執行toy shop demo….
這邊擷取一下cho的翻譯:
[quote]问:我确信如果可能的话,会有人去写一个"wrapper"(让NVIDIA的卡也能运行),在过去不也正是这样吗?

回答:正如(Thorsten和Chris,另外两位在BBS中发言的ATI演示团队工作人员)所言,我们使用了一些在NVIDIA硬件上不支持但是对演示程序能提供非常大帮助的特性(例如R5XX专有的顶点数据格式、3Dc压缩、深度纹理格式、能显著节省内存带宽并提供良好HDR品质的10:10:10:2格式,如果不采用这些数据格式,将无法把庞大的资料塞进显示卡内存里)。

同样,parallax occlusion mapping可以充分发挥我们X1K系列产品的技术优势。

在开发这个技术的时候,我对比了G60、G70和R520在典型的parallax occlusion mapping场景效能,视乎不同的情形。在G60、G70上的运行速度只有R5XX同档次的30%~50%。

如果你想要获得平滑、正确的效果,NVIDIA目前最新的产品显然无法实现。你甚至无法在不使用10:10:10:2、3Dc以及额外顶点数据格式(例如dec3n)的情况下把demo的资料塞进内存中。这个演示使用了庞大的纹理和顶点数据。

当然,有人会想使用变通的执行方式来替代某些我们在这个演示中执行的算法。例如你可以使用relief mapping而不是parallax occlusion mapping。relief mapping技术在ATI和NVIDIA的硬件上都能执行,因为这个技术没有使用到动态分支并且会使用大量的相依性纹理读取。不过在我作的这两种技术的品质比较中,relief mapping如果使用我们目前的数据组是会出现人工化痕迹的。[/quote]

關於3Dc+,就我所知OpenEXR有個2:1的壓縮,不過一來並不確定那是3Dc+,二來也不確定NV4x/G7x可以支援。
其次是10:10:10:2 與 dec3n,這個應該都不是目前NV4x/G7x可以支援的部份。

所以光論這些部分的話,ATI應該不會擔心NV42與 RV530之間的對抗;反過來說NV42也只要撐到G72登場為止就可以了,那應該是06Q1的事情。

話說其實我覺得蠻有趣的地方,就是到時候G72的規模。
目前G72的規模聽說大概210M上下,同樣是5vs12ps,記憶體是128bit,ROP數量不明。
也就是說要是沒什麼結構上的好改進,那7600自己都不見得能夠打贏6800GS….XD
(well,如果G70後續的G7x也能有FAST AF的話,那大概就不成問題了?)

此外,NV42的DDL到底要不要開呢?

附圖:FX3450 aka NV42。

—-
剛好老大手邊有張NV41 on 7800GT’s PCB….拿來玩玩。XD

從2005動畫最萌現況看討論版的水準維持

LH+ thread:
http://www.lovehinaplus.com/phpBB2/viewtopic.php?t=7154

剛好最近出現票數膨脹的外國票….
所以反過來討論到討論版的秩序/品質上的維持部份。

大規模的組織鐵票会使討論串的流向低俗化、

1、一行レス、只投票不po萌文也不支援物資
2、1人多重投票(multi-host)、或是呼朋引伴在短時間内衝高票
3、忽視支援和和討論串流向、只想固票
這些都是最萌的大忌、

正如大部分人所説的、
八雲就算没有海外的組織票一様勝券在握、
但不幸的是海外的票幾乎以上皆是

要切記的是、所謂的「票」同時也是討論的一篇回覆文章

如果某一個普通的討論串7、8成的回覆文章
都是「某大大的説法+1」、「楼主説的好」、
「這個我討厭」、「XX比OO好」之類的回文的話各位作何感想?

(by 村上)

但是反過來看到"彈盡"的狀況,說起來炒熱到這個地步好像也難以不出現品質上的稀釋….強者的資源也不是無窮無盡的。

X1000 series 台灣發表會

今天下午兩點,ATI X1000 series台灣發表會。
聽有去的 ikari 說,是David Wang(Director of Engineering)親自出馬….(wow,ArtX的頭子!XD)

但是好像去的人不多,環視周遭只有十來個….
而且問不到什麼問題。XD

結論來說,ATI 對GPGPU的態度,也是採行會發行API的方式;其他部分(如AVIVO的H.264)則都是未定、觀察。

所以其實乏善可陳嗎….. _A_

source:
http://bbs.gzeasy.com/index.php?showtopic=461982&st=22

Mike houston對R520的一些敘述:
mhouston
GP

Joined: 02 Sep 2003
Posts: 241
Location: Stanford University
Posted: Wed Oct 05, 2005 5:01 pm Post subject: A little R520 info

——————————————————————————–

Now that things are public, I can talk about some things:

The board is 32-bit. The precision on ops is slightly better in general than Nvidia, but not in all cases (from GPUBench precision test). ATI cuts corners, much like Nvidia, when it comes to denorms.

Readback rates are still a problem under GL (450MB/s), but not under DX on Nforce4 or ATI chipsets (900+MB/s). There are performance problems on Intel chipsets for some reason. Still below where I’d like to see them, but at least closer to Nvidia performance.

The board has really good latency hiding, much like the R4XX series. Your performance is generally the max(ALU, tex, branch). Where tex is the total fetch latency: 4 cycles for a 128-bit fetch which is a cache hit, and 8 cycles for a 128-bit streaming fetch. You can look at the ClawHMMer paper for more analysis of latency hiding.

The board supports generallized scatter, yet it’s not currently exposed (no way to do this cleanly in DX, so it might be GL only (<- I’m working on this)…)

The board has 1.5 ALUs. The half can do add/sub but not MUL/MAD/etc. This gives the X1800XT ~120GFlop peak. Raw MAD rate is 83GFlops, which is lower than Nvidia.

Cache bandwidth is 42GB/s and streaming is 21GB/s for the X1800XT.

Branch granularity is ~16 fragments. Branch performance, at least from really basic tests, seems very good.

ATI has claimed to be more committed to supporting academic research and GPGPU in general. They say they will open up a lot more information about their architecture and provide lower level interfaces to access their hardware. Only time will tell how this will play out.

Let me know if you have other questions, and I’ll try to answer them as soon and as well as I can. At the moment, I only have a X1800XL here, so I’ll try to put up some GPUBench results for the board later today on the GPUBench site.

-Mike

Last edited by mhouston on Thu Oct 06, 2005 12:49 am; edited 1 time in total

[quote]The board supports generallized scatter, yet it’s not currently exposed (no way to do this cleanly in DX, so it might be GL only (<- I’m working on this)…)
[/quote]

Posted: Thu Oct 06, 2005 1:17 am Post subject:

——————————————————————————–

You can have an arbitrary number of outputs from a shader, well, I guess the instruction limit, so 512.

You can basically do a[i] = x. The writes are uncached, so there will be a performance penalty (think in the thousands of cycles), but if you do lots of ops, some of the latency can be hidden, at least in theory. You are responsible for making sure fragments don’t clobber each other. Also, you cannot read and write to the same buffer, i.e. no read-modify-write. I haven’t tested it yet, since it’s not exposed currently in any available driver, but the memory controller and memory system were designed to handle this.

Posted: Thu Oct 06, 2005 1:29 am Post subject:

——————————————————————————–

Yes. But, you can also output more than 16 floating point values (4 float4’s) as well. Both are useful. We’ve been asking the graphics card companies for awhile about this one as it solves some issues with variable output from kernels as well as stream filtering. It’s going to be interesting to see if it’s cheaper than the known methods, like Daniel Horn’s chapter in GPU Gems2.

[quote]I just checked the GPUBench page and compared the X1800 to the X800XT PCIe. It seems to me that the computational power remained nearly the same. (instruction issue, scalar vs. vector instruction issue, basic throughput, FP Bandwidth)

The most significant differences seem to be the new branching and 32-Bit support.

So is it faster than the former ATI cards e.g. X850? Or as fast as those cards, but now with 32 Bit support?[/quote]

Posted: Thu Oct 06, 2005 2:07 pm Post subject:

——————————————————————————–

The X1800XL has roughly the same clock rates as the X8XX boards, 500 core/500 mem. The branching, 32-bit, scatter, no dependent texture limit, no dynamic instruction limit, and fully associative cache are the biggest new things. The R520 is ~20% faster than the R4XX clock for clock and has a MUCH better memory subsystem so it handles random reads better. Basically, all our apps got a little faster on the XL, ~10-15%.

The X1800XT has is clocked at 625c/750m, so is substantially faster. We’ve seen compute bound applications get ~30% and memory bound applications get 50-100% depending on the memory access patterns. The later is from the new cache design (many fewer misses) and the memory subsystem handling incoherent reads much better.

從 R520 看 ATI 在 GPGPU 上的優勢

繼上回Radeon X1000系列日本發表會,提出「副卡可以作為PPU」的點子之後,ATI再次對GPGPU這個範疇進行了自我推薦:

http://techreport.com/onearticle.x/8887
Tech-Report報導了ATI與Stanford的Mike Houston合作的這次demo。

要點有二:
1. ATI將與Havok合作,開發物理引擎所需的相關API。
2. R520在GPGPU上的優勢,並且有幾個現成的demo

有一份相關的文件可以參考,並且包含了相關的測試結果。
http://graphics.stanford.edu/~mhouston/public_talks/R520-mhouston.pdf

可以看得出來這些測試裡面X1800XT有相當的優勢。

GROMACS – GPU Implementation:
Written using Brook by non-graphics programmers
– Offloads force calculation to GPU (~80% of CPU time)
– Force calculation on X1800XT is ~3.5X a 3.0GHz P4
– Overall speed up on X1800XT is ~2.5X a 3.0GHz P4
Not yet optimized for X1800XT
– Using ps2b kernels, i.e. no looping
– Not making use of new scatter functionality
The revenge of Ahmdal’s law
– Force calculation no longer bottleneck (38% of runtime)
– Need to also accelerate data structure building (neighbor lists
‧ MUCH easier with scatter support
This looks like a very promising application for GPUs
– Combine CPU and GPU processing for a folding monster!
(from Document)

話說這邊有個有趣的部份,就是Mike Houston倡導的部份:

What GPGPU needs from vendors More information
– Shader ISA
– Latency information
– GPGPU Programming guide (floating point)
‧ How to order code for ALU efficiency
‧ The “real” cost of all instructions
‧ Expected latencies of different types of memory fetches Direct access to the hardware
– GL/DX is not what we want to be using
‧ We don’t need state tracking
‧ Using graphics commands is odd for doing computation
‧ The graphics abstractions aren’t useful for us
– Better memory management Fast transfer to and from GPU
– Non-blocking Consistent graphics drivers
– Some optimizations for games hurt GPGPU performance

What GPGPU needs from the community
Data Parallel programming languages
– Lots of academic research
“GCC” for GPUs
Parallel data structures
More applications
– What will make the average user care about GPGPU?
– What can we make data parallel and run fast?

這段的意思等於是說,GPGPU需要的資訊至少需要和CPU廠商提供出來的資訊同樣詳細;而以過往的經驗來說,這似乎會觸動到NVIDIA的神經….比方說光那個Shader ISA、Latency Information公開就已經觸動到NVIDIA的神經了吧。XD

所以短期內GPGPU這個範疇大概會是ATI比較佔優勢了。

從鋼彈的觀點看なのはA’s

心得:
http://webbbs.gamer.com.tw/readPost.php?brd=GameAC&p=8886&maxpos=9023&thread=-999

捏他圖:
http://mis.im.tku.edu.tw/~fireflyyen19a/new/src/1128820114376.jpg

短評:これ 本当に魔法少女のアニメなの?

其實言下之意是魔法少女アニメ都是低成本動畫。(w

補圖:
http://mis.im.tku.edu.tw/~fireflyyen19a/new/src/1128814999298.jpg

這張GJ XD

====

於是:從魔法少女的觀點看SEED-Destiny
http://webbbs.gamer.com.tw/readPost.php?brd=Gundam&p=8877&rand=20051013

特訓第11天

昨天的AP測試得到成功之後,
早上練完兩個小時之後,就打定主意出來買天線。
剛好又遇上ilo好像早上兩個老師都請假,完全空閒….
於是就順便找出來繞繞。

雖然本來計畫去光華,但是在NOVA拿了5dbi的增益天線、KMall拿了Ultrabay 2000的2nd HDD adapter、順便請他們把電池回收之後,猛然不需要買的、不該買的都弄完了。_A_

於是這邊想到一件事情,就是檢證上回02提到的上好炸豬排。
http://myread02.blogspot.com/2005/09/ilopca.html

店名:添財日本料理(開封店)
店址:台北市開封路1段38號

先引一下02的內文:
[quote]豬排送來,顏色炸得有點深,咬起來相當脆,沒有炸不熟的麵粉感;豬排也相當多汁,他們對豬排的自信可鑑於只淋了些蕃茄醬就給我上桌…XD

勝丼下面的飯淋了某種醬汁,飯本身已經是顆粒分明了,醬汁的味道可食不可言傳。[/quote]

違う。
うまいですが,これはカツ丼ではない
這是純豬排飯,不是丼飯啊啊啊啊啊!
蛋汁和洋蔥都不見了…. _A_
還有,底下白飯真的有淋東西嗎?

實際上這個豬排大概和聯歡小西門、財資味的豬排可以對抗了,比較像是中式的豬排;可是只有豬排和底下的白飯這種配法又很日式….話說豬排的份量蠻不錯的,所以這樣說120元在台北這樣算很棒了;只是順便點的土瓶蒸是敗筆。(死)

—-
回過頭來,把光華逛完一遍之後,我回來發現….我腳起水泡了。
想想不知怎的我講手機大多得走到樓下可能是個原因。(死)
還有加上去的天線看起來好像沒有顯著差異,不過總之繼續放著。
5dbi的天線差不多200元,算是正常價….不過在台中好像整體持有成本可以壓低10%?

晚上繼續。
總之只要沒出去應該都會練滿六小時….
出去的話就是那個時段的兩個小時disable,之後繼續練這樣。
看看這樣持續一個月廢柴度能不能有點起色….(爬走)