{"id":3569,"date":"2009-11-04T00:57:38","date_gmt":"2009-11-04T00:57:38","guid":{"rendered":"http:\/\/192.168.1.4\/wordpress4\/?p=3569"},"modified":"2009-11-04T00:57:38","modified_gmt":"2009-11-04T00:57:38","slug":"w","status":"publish","type":"post","link":"https:\/\/aiplus.idv.tw\/wp\/2009\/11\/04\/w\/","title":{"rendered":"\u4efb\u5929\u5802\u65e9\u5c31\u8cfa\u98fd\u4e86\u5427w"},"content":{"rendered":"<p><a href=\"http:\/\/blog.livedoor.jp\/htmk73\/archives\/621167.html\">http:\/\/blog.livedoor.jp\/htmk73\/archives\/621167.html<\/a><br \/>\n<br \/>\n\u5ca9\u7530\u793e\u9577\u304c\u6c7a\u7b97\u8aac\u660e\u4f1a\u306e\u8cea\u7591\u5fdc\u7b54\u3067\u304a\u307e\u3048\u3089\u3084\u30de\u30b9\u30b3\u30df\u306e\u5831\u9053\u306b\u6012\u308a\u306e\u30b3\u30e1\u30f3\u30c8<\/p>\n<p>\u4ed6\u6839\u672c\u5728\u6f14\u6232\u5427w<br \/>\n<br \/>\n\u60f3\u5230\u9019\u500b\uff1a<\/p>\n<p><a href=\"http:\/\/vipper2ch.blog94.fc2.com\/blog-entry-457.html\">http:\/\/vipper2ch.blog94.fc2.com\/blog-entry-457.html<\/a><br \/>\n<br \/>\n\u4efb\u5929\u5802\u306f\u5f93\u696d\u54e11\u4eba\u8fba\u308a\u3001\u7d0410\u5104\u5186\u306e\u58f2\u4e0a\u9ad8\u3092\u4e0a\u3052\u3066\u3044\u308b\u3089\u3057\u3044<\/p>\n<p>\u73fe\u91d1\u5b58\u91cf\u4e0d\u77e5\u9053\u6709\u591a\u53ef\u6015\u7684\u5730\u6b65w<br \/>\n<br \/>\n\uff08\u96d6\u7136\u4efb\u5929\u5802CM\u7528\u8d85\u5147\u7684\uff09<\/p>\n<p>&#8212;&#8212;-<br \/>\n<br \/>\n<a href=\"http:\/\/forums.nvidia.com\/index.php?showtopic=84440\">http:\/\/forums.nvidia.com\/index.php?showtopic=84440<\/a><br \/>\n<br \/>\nNVIDIA CUDA FAQ version 2.1<\/p>\n<blockquote>\n<p>CUDA FAQ #34:<br \/>\n<br \/>\nIs it possible to run multiple CUDA applications and graphics applications at the same time?<\/p>\n<p>CUDA is a client of the GPU in the same way as the OpenGL and Direct3D drivers are &#8211; it shares the GPU via time slicing. It is possible to run multiple graphics and CUDA applications at the same time, although currently CUDA only switches at the boundaries between kernel executions.<\/p>\n<p>The cost of context switching between CUDA and graphics APIs is roughly the same as switching graphics contexts. This isn&#8217;t something you&#8217;d want to do more than a few times each frame, but is certainly fast enough to make it practical for use in real time graphics applications like games.<\/p>\n<\/blockquote>\n<p>\u54a6\uff1f\u9019\u6a23\u8aaa\u4f86\u5176\u5be6CUDA\u7684application\u548cgraphics\u4e00\u6a23\uff0c\u90fd\u5f97\u9760content switching\u4f86\u5207\u63db&#8230;.\u6240\u4ee5\u8aaaG80~GT200\u90fd\u6c92\u6709\u540c\u6642\u57f7\u884cVS\u3001PS\u7684\u80fd\u529b\uff1f\u9019\u597d\u50cf\u4e0d\u592a\u5c0d&#8230;\u6216\u8005\u8aaa\uff0c\u540c\u6a23\u90fd\u662ftime sharing\u7684\u72c0\u6cc1\u4e0b\uff0cG80~GT200\u90fd\u6709\u9ede\u50cfsingle thread\u7684CPU\u3001\u800cATI\u5f9eXenos\u4ee5\u4f86\u90fd\u4e00\u76f4\u6709mutli thread\u7684\u80fd\u529b&#8230;. \u563f\uff0c\u8003\u616eVLIW\u4e0b\u7684R600\u9084\u662f\u5e38\u670960%\u7684worst case\uff0c\u9019\u6a23\u6211\u9084\u771f\u4e0d\u77e5\u9053\u54ea\u908a\u6548\u7387\u9ad8w<\/p>\n<p>\u4e0d\u904e\u5982\u679c\u60f3\u60f3Fermi\u76f8\u5c0d\u65bcG80~GT200\u670910\u500d\u7684\u6539\u5584\u9019\u9ede\uff0c\u6539\u8b8a\u4e4b\u5f8c\u662f20~25 microsecond(\u00b5s)\u7684\u8a71\uff0c\u90a3\u5c31\u7b97\u662f10\u500d\u4e5f\u5927\u6982\u662f300\u00b5s\u524d\u5f8c\uff0c\u597d\u50cf\u771f\u7684\u4e0d\u6703\u5dee\u5f88\u591a\u3002\u91cd\u9ede\u53ef\u80fd\u9084\u662f\u572816\u500bSM\u90fd\u53ef\u4ee5\u5404\u81ea\u57f7\u884c\u4e0d\u540c\u7684kernel\u9019\u9ede\u4e5f\u8aaa\u4e0d\u5b9a\u3002\u96d6\u7136\u9019\u5c0d\u624b\u65e9\u5c31\u505a\u4e86w<\/p>\n<p><a href=\"http:\/\/www.anandtech.com\/video\/showdoc.aspx?i=3334&amp;p=6\">http:\/\/www.anandtech.com\/video\/showdoc.aspx?i=3334&amp;p=6<\/a><br \/>\n<br \/>\nDerek&#8217;s Conjecture Regarding SP Pipelining and TMT<\/p>\n<blockquote style=\"margin-right:0;\">\n<p>In G80 and GT200, because of the fact that context is stored per warp, even though the SPs are working on an instruction for a different thread in every pipeline stage, they are not working on a different context at every pipeline stage. Each SP processes four threads in a row from the same warp and thus from the same context. Because it is incredibly likely at 1.5GHz that the SPs have more than 4 pipeline stages, we will still see more than one context switch within the pipeline itself, but it still isn&#8217;t down to a different context for every stage.<\/p>\n<\/blockquote>\n<p><a href=\"http:\/\/zergone.blogspot.com\/2009\/10\/fermi-technology-unveiled.html\">http:\/\/zergone.blogspot.com\/2009\/10\/fermi-technology-unveiled.html<\/a><br \/>\n<br \/>\nFermi technology unveiled<\/p>\n<p><a href=\"http:\/\/techreport.com\/articles.x\/17670\/2\">http:\/\/techreport.com\/articles.x\/17670\/2<\/a><br \/>\n<br \/>\nBetter scheduling, faster switching<\/p>\n<blockquote style=\"margin-right:0;\">\n<p>&#8220;Fermi avoids this inefficiency by executing up to 16 different kernels concurrently, including multiple kernels on the same SM. The limitation here is that the different kernels must come from the same CUDA context-so the GPU could process, say, multiple PhysX solvers at once, if needed, but it could not intermix PhysX with OpenCL.&#8221;<\/p>\n<p>\u770b\u8d77\u4f86\u4e3b\u8981\u7684\u9650\u5236\u5c31\u662f\u4e0d\u80fd\u5920\u597d\u5e7e\u500b\u4e0d\u540c\u7684\u7a0b\u5f0f\u540c\u6642\u5229\u7528GPU&#8230;.\u4e0d\u904e\u672c\u4f86\u7684\u8a71\u540c\u4e00\u500bapplication\u88e1\u9762\u540c\u6642\u6709graphic\u548cCUDA\u4f3c\u4e4e\u4e5f\u4e0d\u6703\u6709\u9019\u500b\u554f\u984c\uff0c\u53ea\u662fFermi\u6548\u7387\u61c9\u8a72\u6703\u66f4\u9ad8\u9ede\u3002<\/p>\n<p>To tackle that latter sort of problem, Fermi has much faster context switching, as well. Nvidia claims context switching is ten times the speed it was on GT200, as low as 10 to 20 microseconds. Among other things, intermingling GPU computing with graphics ought to be much faster as a result. (Incidentally, AMD tells us its Cypress chip can also run multiple kernels concurrently on its different SIMDs. In fact, different kernels can be interleaved on one SIMD.)<\/p>\n<\/blockquote>\n<p>&#8212;&#8211;<br \/>\n<br \/>\n<a href=\"http:\/\/www.intrinsity.com\/index.php\/articles\/64-hot-rodding\">http:\/\/www.intrinsity.com\/index.php\/articles\/64-hot-rodding<\/a><br \/>\n<br \/>\nHot-Rodding the Cortex-A8<\/p>\n<p>Because Fast14 logic gates are 25% to 50% faster than static logic gates, the processor can do more work per clock cycle without altering the basic design of the instruction pipelines and functional blocks. Fast14 is particularly efficient for muxes and other elements with wide structures. Intrinsity also uses optimized static logic, custom circuits, and standard cells. (See MPR 8\/13\/01-02, &#8220;Intrinsity&#8217;s Dynamic Designs.&#8221;) Figure 1 shows Intrinsity&#8217;s design flow.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>http:\/\/blog.livedoor.jp\/htmk73\/archives\/621167.html \u5ca9\u7530\u793e &hellip; <a href=\"https:\/\/aiplus.idv.tw\/wp\/2009\/11\/04\/w\/\" class=\"more-link\">\u95b1\u8b80\u5168\u6587 <span class=\"screen-reader-text\">\u4efb\u5929\u5802\u65e9\u5c31\u8cfa\u98fd\u4e86\u5427w<\/span> <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[],"class_list":["post-3569","post","type-post","status-publish","format-standard","hentry","category-cell"],"_links":{"self":[{"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/posts\/3569","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/comments?post=3569"}],"version-history":[{"count":0,"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/posts\/3569\/revisions"}],"wp:attachment":[{"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/media?parent=3569"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/categories?post=3569"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/aiplus.idv.tw\/wp\/wp-json\/wp\/v2\/tags?post=3569"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}